cd /news/artificial-intelligence/accuracy-without-grounding-diagnosin… · home topics artificial-intelligence article
[ARTICLE · art-61428] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Accuracy Without Grounding: Diagnosing Visual Dependency Dissociation in Video LLM Benchmarks

A new study from researchers at arXiv introduces the Visual Dependency Gap (VDG) metric to audit whether video large language models (LLMs) rely on visual information for correct answers. Testing twenty models from 2B to 78B parameters on MVBench, the authors found that accuracy and visual dependency are separable: models differed on original video (p = 0.0003) but not on black screens (p = 0.53), and temporal order contributed near-zero accuracy improvement. The findings suggest that benchmark accuracy may not reflect genuine visual understanding, motivating VDG as a standard audit metric.

read1 min views41 publishedJul 16, 2026

arXiv:2607.13305v1 Announce Type: new Abstract: Benchmark accuracy in video large language models (LLMs) is often treated as evidence of visual understanding. We audit this assumption across twenty models spanning 2-78B parameters and ten architecture families. We introduce the Visual Dependency Gap (VDG), the difference in per-question correctness between original-video and black-screen conditions. Paired McNemar tests on MVBench show that accuracy and visual dependency are separable: models differ on original video (p = 0.0003) but not on black screens (p = 0.53). Across models, task-type rankings are stable: Attribute Perception is strongly visual, whereas Temporal Reasoning approaches the language-only baseline. A diagnostic ladder from black screen to single frame, shuffled frames, and original video reveals that frame diversity supplies most of the visual benefit, while temporal order contributes near-zero accuracy across sixteen open-weight models. An ablation from 0.5 to 24 FPS rules out sparse sampling as the cause. H.264 experiments further show that stable aggregate accuracy conceals bidirectional question-level answer flips. The diagnostic also generalizes to four API-accessed models, whose VDG values range from 0.025 to 0.315. These results motivate VDG as a standard audit for whether video benchmarks measure visually grounded capability. Code is available at https://github.com/JaeLee18/accuracy-without-grounding.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @mvbench 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/accuracy-without-gro…] indexed:0 read:1min 2026-07-16 ·