{"slug": "an-exam-for-active-observers", "title": "An Exam for Active Observers", "summary": "A new benchmark called ActiveVision reveals that today's multimodal large language models (MLLMs) lack robust active visual observation, with the highest-scoring model, GPT-5.5 at its highest reasoning-effort tier, solving only 10.6% of items and scoring zero on 11 of 17 tasks, while three human participants averaged 96.1%. The benchmark, introduced by researchers in a paper on arXiv, comprises 17 tasks across 3 categories designed to force repeated visual perception rather than a single static description, and even Claude Fable 5, despite topping reasoning and coding leaderboards, solved just 3.5%.", "body_md": "# Computer Science > Computer Vision and Pattern Recognition\n\n[Submitted on 17 Jul 2026]\n\n# Title:An Exam for Active Observers\n\n[View PDF](/pdf/2607.16165)\n\n[HTML (experimental)](https://arxiv.org/html/2607.16165v1)\n\nAbstract:Human vision is a closed loop: gaze is continuously redirected by intermediate hypotheses rather than a single snapshot. Decades of psychophysics and cognitive science have argued that this active observation is essential for a wide range of tasks. Whether today's multimodal large language models (MLLMs) exercise active observation is an empirical question that current vision-language benchmarks do not answer. We introduce ActiveVision, a benchmark that makes active observation measurable for MLLMs, comprising 17 tasks across 3 categories. Tasks are designed to force repeated visual perception rather than a single static description. Frontier MLLMs collapse on ActiveVision: the highest-scoring model we evaluate, GPT-5.5 at the highest exposed reasoning-effort tier, solves only 10.6% of items and scores zero on 11 of the 17 tasks, and even Claude Fable 5, despite topping most reasoning and coding leaderboards, solves just 3.5%, far behind three human participants who average 96.1%. Furthermore, much of the gap persists even when models write and run their own vision code: such code is unreliable on realistic imagery, and catching its failures itself requires the active perception the models lack. Together, these results indicate that current MLLMs lack robust active visual observation, motivating architectures and training objectives that close the perception-reasoning loop.\n\n### Current browse context:\n\ncs.CV\n\n### References & Citations\n\nLoading...\n\n# Bibliographic and Citation Tools\n\nBibliographic Explorer\n\n*(*[What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))\nConnected Papers\n\n*(*[What is Connected Papers?](https://www.connectedpapers.com/about))\nLitmaps\n\n*(*[What is Litmaps?](https://www.litmaps.co/))\nscite Smart Citations\n\n*(*[What are Smart Citations?](https://www.scite.ai/))# Code, Data and Media Associated with this Article\n\nalphaXiv\n\n*(*[What is alphaXiv?](https://alphaxiv.org/))\nCatalyzeX Code Finder for Papers\n\n*(*[What is CatalyzeX?](https://www.catalyzex.com))\nDagsHub\n\n*(*[What is DagsHub?](https://dagshub.com/))\nGotit.pub\n\n*(*[What is GotitPub?](http://gotit.pub/faq))\nHugging Face\n\n*(*[What is Huggingface?](https://huggingface.co/huggingface))\nScienceCast\n\n*(*[What is ScienceCast?](https://sciencecast.org/welcome))# Demos\n\n# Recommenders and Search Tools\n\nInfluence Flower\n\n*(*[What are Influence Flowers?](https://influencemap.cmlab.dev/))\nCORE Recommender\n\n*(*[What is CORE?](https://core.ac.uk/services/recommender))# arXivLabs: experimental projects with community collaborators\n\narXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.\n\nBoth individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.\n\nHave an idea for a project that will add value for arXiv's community? [ Learn more about arXivLabs](https://info.arxiv.org/labs/index.html).", "url": "https://wpnews.pro/news/an-exam-for-active-observers", "canonical_source": "https://arxiv.org/abs/2607.16165", "published_at": "2026-07-24 16:37:00+00:00", "updated_at": "2026-07-24 16:52:32.556723+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "computer-vision", "ai-research"], "entities": ["ActiveVision", "GPT-5.5", "Claude Fable 5", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/an-exam-for-active-observers", "markdown": "https://wpnews.pro/news/an-exam-for-active-observers.md", "text": "https://wpnews.pro/news/an-exam-for-active-observers.txt", "jsonld": "https://wpnews.pro/news/an-exam-for-active-observers.jsonld"}}