The open HSS benchmark tests inference across 288 images and 234 videos; 20 human participants scored 93.1% under the same free-form conditions.
By [Ryan Merket](https://runtimewire.com/author/ryan-merket)
· Published
Primary source: [X](https://x.com/ScaleAILabs/status/2107920376756117903)
Why it matters #
HSS measures visual inference that standard perception and knowledge tests can miss, while its results suggest extra reasoning time alone may not correct models' failures to identify cues, spatial relationships and social context.
Scale AI and Elorian released Humanity's Sixth Sense (HSS) on October 7th, 2026, a benchmark that puts a stubborn weakness in multimodal AI on a scorecard: the top model tested answered correctly on 53.6% of tasks, against 93.1% for human participants. Scale's announcement names GPT-6-astra as the leading model at launch; the median of 25 models was 30.9%.
Elorian co-founder and CEO Andrew Dai helped build the benchmark with Scale. Before starting Elorian, Dai worked at Google Brain and DeepMind, where the company's team page credits him with leading Gemini's data area and co-leading pretraining for PaLM 2. The partnership puts Elorian's visual-thinking thesis to a public test: its site argues that current vision-language models translate images into language before reasoning, a process it considers fragile for spatial and relational tasks.
HSS contains 522 open-ended tasks based on 288 images and 234 video clips, totaling 17.6 hours. Its prompts ask models to infer events and relationships that are not fully spelled out in the frame: what caused an action, whether a vehicle can fit in a parking space, why someone slowed down after running, or what a partially obscured pattern implies. The benchmark and leaderboard divide the questions into temporal and causal dynamics, physical and spatial logic, social understanding, and abstract and contextual inference.
The design aims to separate visual inference from factual recall and multiple-choice elimination. Models respond in free form; an automated language-model judge grades each answer against criteria written for the task. Scale says 3,466 tasks were authored and 522 survived three independent review rounds. The company reports that five models regraded with judges from three vendors kept the same rankings, with more than 95% agreement. The dataset is available on Hugging Face under an MIT license.
The headline score is a benchmark result, not a measure of how often a model will fail in everyday use. Scale's leaderboard defines its main metric, pass@1, as the average across three attempts per task; the human comparison comes from 20 participants answering under free-form conditions. The test itself is also small relative to the range of visual situations that people encounter. Scale notes that some source media comes from the public web, so prior exposure in model training cannot be ruled out. For questions about people's beliefs or intentions, the score measures agreement with annotator judgment, not an objective reading of the people shown.
The errors point toward perception and inference as the bottleneck. Across 8,573 model failures, Scale attributes 94% to perception or latent inference and 5% to faulty logic. Social-understanding tasks were the weakest area for 21 of the 25 models. Video was harder than still images for 23 models, with a 7.3-point average drop. Those results complicate the idea that giving a model more time to reason will fix visual mistakes: Scale says models used an average of about 4,000 reasoning tokens per task, while greater effort reduced GPT-6-astra's score on some subdomains.
Tool use helped, within limits. Scale tested models inside Claude Code and Codex, where they could crop, zoom, search the web and sample media again. The best setup scored 59.3% on a 388-task subset, a different test slice from the full benchmark's 53.6% result. Closer inspection reduced errors such as missed visual cues, while errors interpreting a two-dimensional overlap as three-dimensional alignment barely changed, falling from 102 to 101 in Scale's analysis.
That gap is the practical point for Elorian, which is building AI around visual reasoning rather than treating vision as an input to text-based analysis. The company's team page also lists co-founder Yinfei Yang, whose prior work includes multimodal research at Apple and Google. Elorian says it has raised $55 million, with Striker Ventures, Menlo Ventures, Altimeter, NVIDIA and 49 Palms among its backers. HSS gives labs a shared test for one part of that thesis; its small task set and human-judged criteria leave the harder question of how well benchmark gains transfer to real-world environments unanswered.