cd /news/artificial-intelligence/dochop-benchmarking-out-of-domain-mu… · home topics artificial-intelligence article
[ARTICLE · art-119785] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense Documents

Researchers introduced DocHop, a benchmark for integrated chart-context reasoning in document-style images, finding that the best multimodal large language model achieves only 62.83% accuracy compared to over 90% for human annotators across 2,074 examples. The benchmark, built via a stochastic logic-first generation pipeline with controllable reasoning depth and visual density, covers six task categories and reveals performance degradation as reasoning complexity increases.

read1 min views1 publishedSep 3, 2026

arXiv:2609.02059v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved strong performance on structured visual understanding tasks such as chart and document question answering. However, existing benchmarks typically evaluate these domains in isolation, leaving underexplored a key capability: whether models can use textual context to determine how chart evidence should be selected, interpreted, and aggregated. We introduce DocHop, a benchmark for integrated chart--context reasoning in document-style images. In DocHop, the document narrative specifies multi-step compositional constraints, while charts provide the corresponding data values. Questions are grounded on a semantic reference label defined in the narrative, requiring models to resolve target entities from context before aggregating evidence across multiple charts. To enable systematic evaluation, we construct DocHop via a stochastic logic-first generation pipeline with controllable reasoning depth and visual density, covering 2,074 examples across six task categories. Experiments on a wide range of proprietary and open-source MLLMs show a substantial gap to human performance: annotators achieve over 90% accuracy, while the best model reaches only 62.83%. Reasoning-enhanced models consistently show improved results, but performance degrades as reasoning complexity increases. Overall, DocHop provides a controlled testbed for challenging multi-hop document reasoning.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @dochop 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/dochop-benchmarking-…] indexed:0 read:1min 2026-09-03 ·