cd /news/computer-vision/testhallvqa-exploring-lvlms-document… · home topics computer-vision article
[ARTICLE · art-129812] src=arxiv.org ↗ pub= topic=computer-vision verified=true sentiment=· neutral

TestHallVQA: Exploring LVLMs' Document-Level Reasoning under Redundant Contexts from Scientific Exams

Researchers introduced TestHallVQA, a multi-image visual question answering benchmark built from scientific exams that combines document-level scale with human-examination difficulty, according to the arXiv paper 2609.13158v1. The team also proposed a new metric, F1-R², which jointly quantifies large vision-language models' computational reasoning capability and their evidence retrieval robustness against document-level redundancy. Experiments on mainstream LVLMs revealed latent deficiencies across multiple dimensions, with datasets, code, and theoretical derivations available at the project's GitHub repository.

by read1 min views1 publishedSep 15, 2026

arXiv:2609.13158v1 Announce Type: new Abstract: Large Vision--Language Models (LVLMs) are increasingly expected to perform visual question answering (VQA) over planar media. However, existing planar VQA benchmarks typically emphasize isolated challenges: some emphasize long-document understanding with limited reasoning depth, while others require complex visual reasoning but remain restricted to single-page, noise-free settings. Moreover, through theoretical analysis, we identify the impact of irrelevant visual tokens, which leads to measurable performance degradation but has received little attention with respect to systematic quantification. To address these limitations, we introduce TestHallVQA, a multi-image VQA benchmark that simultaneously embodies document-level scale and the difficulty of human examinations, while providing comprehensive task coverage. Leveraging TestHallVQA's ability to controllably inject multi-level contextual redundancy, we further propose a novel metric, F1-R\textsuperscript{2}, which jointly quantifies LVLMs' computational reasoning capability and their evidence retrieval robustness against document-level redundancy. Extensive experiments and analyses on mainstream LVLMs reveal their latent deficiencies across multiple dimensions, offering concrete insights and directions for future research. The associated datasets, code, and complete theoretical derivations are available at https://github.com/yqyu2317/TestHallVQA-benchmark.

── more in #computer-vision 4 stories · sorted by recency
── more on @testhallvqa 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/testhallvqa-explorin…] indexed:0 read:1min 2026-09-15 ·