16:37
2026-07-24
arxiv.org
artificial-intelligence
An Exam for Active Observers
A new benchmark called ActiveVision reveals that today's multimodal large language models (MLLMs) lack robust active visual observation, with the highest-scoring model, GPT-5.5 at its highest reasoninβ¦