cd /news/artificial-intelligence/beyond-accuracy-auditing-spatial-pro… · home topics artificial-intelligence article
[ARTICLE · art-85585] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Beyond Accuracy: Auditing Spatial Provenance in Visual Token Pruning for OCR-Critical MLLM Inference

A new arXiv study (2608.00077v1) audits visual-token pruning in OCR-critical multimodal large language models (MLLMs), finding that accuracy alone misses failures where correct answers lack local token provenance. On locked image-disjoint confirmation, Qwen Target at 30% retention achieved 0.786 accuracy versus 0.783 for Full (paired difference +0.003, 95% CI [-0.014, +0.020]), yet same-budget Target, Random, and Grid retained positive-support coverage of 0.620, 0.270, and 0.318. Across Qwen3-VL-8B, LLaVA-1.5-7B, and InternVL3.5-8B, materialized prefixes yielded up to 4.32x batch-prefill speedup and 76.4% lower incremental peak memory, but full-validation TextVQA and DocVQA showed favorable target-verification points do not imply task-general compression.

read1 min views1 publishedAug 4, 2026

arXiv:2608.00077v1 Announce Type: new Abstract: Visual-token pruning is usually judged by answer quality at a fixed retention budget. For text-rich multimodal large language models (MLLMs), this protocol can miss a distinct failure: an answer remains correct even when no retained token is locally traceable to the small OCR region that supports it. We turn this blind spot into an evidence-risk audit that couples answer behavior with geometric token-origin provenance, interventions, and realized cost; transparent training-free selectors isolate controlled operating points. On locked image-disjoint confirmation, Qwen Target at 30% retention has observed accuracy 0.786 versus 0.783 for Full (paired image-cluster difference +0.003, 95% CI [-0.014, +0.020]), yet same-budget Target, Random, and Grid retain sharply different positive-support coverage: 0.620, 0.270, and 0.318. Across Qwen3-VL-8B, LLaVA-1.5-7B, and InternVL3.5-8B, matched controls, interventions, detector tests, and external methods reveal model-specific quality-risk-traceability frontiers that accuracy alone does not expose. Materialized prefixes yield up to 4.32x batch-prefill speedup and 76.4% lower incremental peak memory; full-validation TextVQA and DocVQA further show that favorable target-verification points do not imply task-general compression. Visual-token pruning should therefore report surviving spatial provenance and realized cost alongside quality and compression.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/beyond-accuracy-audi…] indexed:0 read:1min 2026-08-04 ·