HSS: AI Can Read the Books. Can It See the Gap?
Try the bookshelf question before reading the answer. HSS exposes a gap in AI visual reasoning—and a product choice: when should an agent look harder, and when should it read the underlying state?

Could you fit two more books on this shelf?
Look at the row closest to the camera, with the white, orange and blue volumes. The question is whether there’s room in front of the light-blue book, before the divider, for two more books about the size of the white volumes already there. Make your call before scrolling.

The reference answer is yes: the available gaps are enough.
In the versions and settings tested, GPT-6-astra, Gemini-3.8-flash and Claude-Opus-5.5 each answered three times. All nine answers said no. This is a published example from Elorian, which developed the Humanity’s Sixth Sense, or HSS, benchmark with Scale AI.
That failure interests me more than another impressive coding score. Sending an AI a screenshot comes with an easy assumption: it has seen what I see, and now we can discuss what to do. If that first step goes wrong, a beautifully organized explanation can make the mistake harder to notice.
What the HSS benchmark actually measures
The name sounds mystical. The questions are ordinary: Is there enough space? Is a chain under tension? Where did someone in a video come from?
Recognizing books gets you only partway through the shelf question. You also need to judge the gaps, the angles and the space available if the volumes are straightened. HSS tests these relationships: what you can infer from a scene beyond naming the objects in it.
Scale’s launch post describes 522 open-ended tasks, comprising 288 images and 234 videos. They cover spatial reasoning, time and causality, social understanding and contextual inference. Models have to produce their own answers rather than choose from a list.
At release, the human baseline was 93.1%. The best tested model, GPT-6-astra at maximum reasoning effort, scored 53.6%. That’s a gap of 39.5 percentage points.

The human baseline came from 20 participants. The tasks also went through a selection process designed to retain questions on which people could reach agreement. This is a substantial gap on a particular kind of visual judgment; it isn’t a universal intelligence score.
One detail matters when reading the leaderboard methodology: pass@1 is averaged over three attempts per task. A model doesn’t get full credit simply because one of its three answers happens to be right. And the bookshelf result records specific model versions under specific conditions. It doesn’t mean every future chat will miss that question.
Check what it saw before debugging what it thought
The paper’s analysis of 8,573 failed responses classified roughly 53% as perception errors, 41% as failures to infer implicit information and 5% as logical reasoning errors. The method is in Appendix D.
Those labels need a qualification. They were assigned by another model, Claude-Opus-5, using the question, reference answer and response. They weren’t a case-by-case human diagnosis, and the classifier wasn’t looking at the original image. So this isn’t evidence that AI has solved logic and merely needs better eyesight.
It does suggest a useful troubleshooting habit: ask the assistant to identify the evidence it used before asking it to reason harder.
Consider a browser agent that mistakes a disabled button for an active one, or a picture of a control for a control it can click. A more elaborate plan still starts from a false premise. That’s my application of the result; HSS didn’t directly test browser operation.
“Think again” is a cheap instruction to give. It can be an expensive substitute for finding out what information is missing.
A better view can help. An absent view is another problem.
The researchers also tested tools. On the same 388-task subset, GPT-6-astra improved from 54.4% with direct answers to 59.3% in a Codex tool environment: a gain of 4.9 percentage points. Appendix F gives the comparison. Subtracting the full benchmark’s 53.6% would mix two different sets of tasks.
Zooming in might reveal a gap the model overlooked. It won’t reveal the back of an object that was never photographed.
If I’m asking whether a cabinet will fit through a doorway, I want the assistant to distinguish a blurry doorframe from an unknown cabinet depth. The first problem might need a closer photo. The second needs a measurement. Burning another round of reasoning on both is an expensive way to avoid asking the right question.
I’d like assistants to get better at proposing one additional observation that could change the answer. A request for a side view can be more useful than a page of analysis, provided the assistant explains which uncertainty that view would resolve.
The product fix I’d start with: give agents readable state
For software, some of that uncertainty is optional. A button has a disabled state. A published article has a database record and a public URL. A completed transaction has a server response. A screenshot shows a presentation of those facts, with some information missing.
When an assistant works on my site, I’d like each important action to come with something I can check: which control it identified, what state it read and what confirmed the result afterward.
That makes exposing reliable state a product priority for me. In an application that already has an API or accessible controls, asking an agent to repeatedly reconstruct the same facts from pixels adds a problem we could have avoided.
Visual reasoning will still matter, especially outside tidy software environments. But I don’t need a model to become an expert at guessing whether an article was published when the publishing system can tell it.
This changes the handoff panel I proposed in my Claude Mods article. Alongside the work, it should let me inspect the page state, image region or video timestamp used to justify a decision. That’s a design I’d want to build, not a feature I’m claiming to have tested here.
The bookshelf question catches a blind spot in how I use AI: I tend to challenge the explanation after accepting its description of the scene.
Next time, I’ll start one step earlier: “Show me where you see that.”
Scope note: the dataset documentation reports 3,466 initial tasks, with 522 retained after three review rounds. The failure analysis covers 24 models; the main results include a later-added 25th. Scores here refer to the release reported in the paper, not a live ranking. Error-category percentages are rounded.