{"slug": "beyond-accuracy-auditing-spatial-provenance-in-visual-token-pruning-for-ocr-mllm", "title": "Beyond Accuracy: Auditing Spatial Provenance in Visual Token Pruning for OCR-Critical MLLM Inference", "summary": "A new arXiv study (2608.00077v1) audits visual-token pruning in OCR-critical multimodal large language models (MLLMs), finding that accuracy alone misses failures where correct answers lack local token provenance. On locked image-disjoint confirmation, Qwen Target at 30% retention achieved 0.786 accuracy versus 0.783 for Full (paired difference +0.003, 95% CI [-0.014, +0.020]), yet same-budget Target, Random, and Grid retained positive-support coverage of 0.620, 0.270, and 0.318. Across Qwen3-VL-8B, LLaVA-1.5-7B, and InternVL3.5-8B, materialized prefixes yielded up to 4.32x batch-prefill speedup and 76.4% lower incremental peak memory, but full-validation TextVQA and DocVQA showed favorable target-verification points do not imply task-general compression.", "body_md": "arXiv:2608.00077v1 Announce Type: new\nAbstract: Visual-token pruning is usually judged by answer quality at a fixed retention budget. For text-rich multimodal large language models (MLLMs), this protocol can miss a distinct failure: an answer remains correct even when no retained token is locally traceable to the small OCR region that supports it. We turn this blind spot into an evidence-risk audit that couples answer behavior with geometric token-origin provenance, interventions, and realized cost; transparent training-free selectors isolate controlled operating points. On locked image-disjoint confirmation, Qwen Target at 30% retention has observed accuracy 0.786 versus 0.783 for Full (paired image-cluster difference +0.003, 95% CI [-0.014, +0.020]), yet same-budget Target, Random, and Grid retain sharply different positive-support coverage: 0.620, 0.270, and 0.318. Across Qwen3-VL-8B, LLaVA-1.5-7B, and InternVL3.5-8B, matched controls, interventions, detector tests, and external methods reveal model-specific quality-risk-traceability frontiers that accuracy alone does not expose. Materialized prefixes yield up to 4.32x batch-prefill speedup and 76.4% lower incremental peak memory; full-validation TextVQA and DocVQA further show that favorable target-verification points do not imply task-general compression. Visual-token pruning should therefore report surviving spatial provenance and realized cost alongside quality and compression.", "url": "https://wpnews.pro/news/beyond-accuracy-auditing-spatial-provenance-in-visual-token-pruning-for-ocr-mllm", "canonical_source": "https://arxiv.org/abs/2608.00077", "published_at": "2026-08-04 04:00:00+00:00", "updated_at": "2026-08-04 04:38:45.502957+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "computer-vision", "ai-research"], "entities": ["arXiv", "Qwen", "Qwen3-VL-8B", "LLaVA-1.5-7B", "InternVL3.5-8B", "TextVQA", "DocVQA"], "alternates": {"html": "https://wpnews.pro/news/beyond-accuracy-auditing-spatial-provenance-in-visual-token-pruning-for-ocr-mllm", "markdown": "https://wpnews.pro/news/beyond-accuracy-auditing-spatial-provenance-in-visual-token-pruning-for-ocr-mllm.md", "text": "https://wpnews.pro/news/beyond-accuracy-auditing-spatial-provenance-in-visual-token-pruning-for-ocr-mllm.txt", "jsonld": "https://wpnews.pro/news/beyond-accuracy-auditing-spatial-provenance-in-visual-token-pruning-for-ocr-mllm.jsonld"}}