cd /news/artificial-intelligence/perceptionbench-evaluating-atomic-vi… · home topics artificial-intelligence article
[ARTICLE · art-78031] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models

Researchers introduce PerceptionBench, a benchmark designed to evaluate atomic visual perception in multimodal large language models (MLLMs), finding that no model reaches 60% accuracy across ten atomic perceptual capabilities. The benchmark, based on an error taxonomy from 42 existing benchmarks, reveals that perception-related hallucination is the weakest capability on average and that similar overall scores conceal divergent capability profiles.

read1 min views1 publishedJul 29, 2026

arXiv:2607.24957v1 Announce Type: new Abstract: We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmarks often fail to isolate perception: holistic evaluations conflate perceptual errors with failures in reasoning or domain knowledge, while application-driven benchmarks only cover narrow, fragmented domains shaped by heuristic designs. To address these limitations, PerceptionBench adopts a bottom-up approach: by diagnosing the earliest failure points in the responses of frontier MLLMs across 42 existing benchmarks, we construct an error taxonomy whose perception branch defines ten atomic perceptual capabilities. Guided by this taxonomy, we construct 3,000 verified questions with short, unambiguous answers, each isolating a single capability, with difficulty stemming from perception rather than reasoning or knowledge. Benchmark results across sixteen frontier MLLMs reveal that atomic perception remains largely unsolved---no model reaches 60% accuracy, perception-related hallucination is the weakest capability on average, and similar overall scores conceal sharply divergent capability profiles. PerceptionBench thus provides a capability-level standard for measuring and diagnosing the visual perception boundaries of MLLMs.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @perceptionbench 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/perceptionbench-eval…] indexed:0 read:1min 2026-07-29 ·