AI models fail a basic human intuition test that humans ace with ease Scale AI and research group Elorian released a benchmark on October 7 called Humanity's Sixth Sense, on which humans scored 93.1% while the best of 25 tested models, OpenAI's GPT-6-astra, reached only 53.6% at its maximum reasoning setting and the median model scored 30.9%. The benchmark comprises 522 open-ended tasks drawn from 288 images and 234 video clips totaling 17.6 hours of footage, probing spatial layout, cause and effect, unwritten social rules and abstract patterns. Social understanding was the worst domain for 21 of the 25 models, which averaged 24.4% on social tasks versus 34.1% across the other three domains combined, a gap Scale AI frames as central to robotics and AI assistant deployment. Scale AI's newest benchmark asks AI models to read a room, not solve an equation, and nearly all of them fail badly. Humans score 93.1% on tasks about social cues and physical common sense. The best model manages 53.6%. Forget the math Olympiad headlines for a second. Scale AI and the research group Elorian released a benchmark on October 7 called Humanity's Sixth Sense, and it measures something far more mundane than frontier math: whether an AI can look at a photo or a short video clip and infer what any person would infer instantly, who's about to get hurt, who's lying, what happens next. Across 25 models tested, the median score was 30.9%. Humans hit 93.1%. The benchmark is built from 522 open-ended tasks spanning 288 images and 234 video clips, 17.6 hours of footage in total, according to Scale's own writeup of the research. Each item comes with a human-written prompt that probes something people infer at a glance: the spatial layout of a scene, a likely cause and effect, an unwritten social rule, or an abstract pattern. No model comes close to matching a human on any of it. The strongest performer, OpenAI's GPT-6-astra, reached 53.6% even running at its maximum reasoning setting. That's not a rounding error behind humans. It's a 40-point gap, on tasks most adults would call easy. Break the results down by category and the picture gets more specific. Social understanding, tasks about theory of mind, emotional state, and who defers to whom in a given scene, was the single worst domain for 21 of the 25 models Scale tested. Models averaged just 24.4% accuracy on social tasks, against 34.1% across the other three domains combined. A model can increasingly solve a proof. It still struggles to tell you why two people in a photo look uncomfortable standing next to each other. Scale AI names Francis deSouza as CEO to drive its enterprise and government AI push https://startupfortune.com/scale-ai-names-francis-desouza-as-ceo-to-drive-its-enterprise-and-government-ai-push/ Scale AI has named Francis deSouza, former Google Cloud COO and ex-CEO of Illumina, as its permanent chief executive, effective August 10. He takes over from interim CEO Jason Droege, who held the role after founder Alexandr Wang departed for Meta as part of the $14.3 billion deal that gave Meta a 49% stake in Scale. DeSouza's cloud and enterprise... - Scale AI names new CEO Francis https://startupfortune.com/scale-ai-names-francis-desouza-as-ceo-to-drive-its-enterprise-and-government-ai-push/ - enterprise AI data labeling company leadership https://startupfortune.com/scale-ai-names-francis-desouza-as-ceo-to-drive-its-enterprise-and-government-ai-push/ That gap matters more than it might sound. Spatial reasoning, causal inference, and social understanding aren't academic curiosities, they're the exact skills a robot needs to not bump into a person, or an AI assistant needs to not misread a frustrated customer's tone. Frontier labs have spent 2026 racing to post record scores on math and coding benchmarks. GPT-6-astra itself reportedly cleared 98% on FrontierMath Tier 4 and posted near-perfect results on ARC-AGI-3 interactive reasoning tasks, according to OpenAI's own release notes. Those numbers made headlines. Humanity's Sixth Sense is a blunt reminder that the same model falls apart on the kind of judgment a toddler develops before it can talk. Scale AI isn't new to this exercise. The company built Humanity's Last Exam with the Center for AI Safety, a benchmark of roughly 2,500 expert-level questions that top models have slowly chewed through, climbing from single digits at its early 2025 launch to over 53% by mid-2026. That benchmark was designed to resist saturation on abstract, text-based reasoning. Humanity's Sixth Sense does something different: it tests whether models have caught up on the kind of reasoning nobody has to study for. They haven't. And the honest reading of these numbers isn't that today's models are dumb, it's that intelligence isn't one thing. A system can out-prove a mathematician and still misjudge a facial expression a six-year-old would read correctly. Scale's benchmark gives that gap a number for the first time, and it's a bigger number than most people tracking AI capability claims would have guessed. For anyone building AI agents meant to operate in the physical world, rather than just a chat window, that's the number worth watching. A coding assistant that aces a leaderboard doesn't need a sixth sense. A robot in a warehouse, or an assistant meant to notice a user is upset, does. Right now, on Scale's own evidence, none of them have it. Also read: Scott Aaronson Says AI Labs Are Quietly Testing Models Against Encryption https://startupfortune.com/scott-aaronson-says-ai-labs-are-quietly-testing-models-against-encryption/ • Samsung's LittleBit squeezes AI models to a tenth of a bit per weight https://startupfortune.com/samsungs-littlebit-squeezes-ai-models-to-a-tenth-of-a-bit-per-weight/ • Researchers find letting AI coding agents write their own tests backfires https://startupfortune.com/researchers-find-letting-ai-coding-agents-write-their-own-tests-backfires/ This article is posted in AI News https://startupfortune.com/category/ai/ , check it out for more related stories. OpenAI Chief Scientist Jakub Pachocki Warns No AI Lab Is Ready to Scale Safely https://startupfortune.com/openai-chief-scientist-jakub-pachocki-warns-no-ai-lab-is-ready-to-scale-safely/ OpenAI chief scientist Jakub Pachocki published an essay saying no AI lab, including OpenAI, has solved alignment well enough to keep scaling at full speed, just three days after OpenAI shipped GPT-6 Astra, its first model to cross a Critical cybersecurity threshold. Sam Altman called the essay 'an important post' even as Astra keeps rolling out... - how to safely scale autonomous AI systems https://startupfortune.com/openai-chief-scientist-jakub-pachocki-warns-no-ai-lab-is-ready-to-scale-safely/ - AI lab safety concerns for frontier models https://startupfortune.com/openai-chief-scientist-jakub-pachocki-warns-no-ai-lab-is-ready-to-scale-safely/ Join the discussion Open in the community → https://startupfortune.com/community/ Almost there. Sign in and your reply posts straight away.