{"slug": "beyond-blind-compliance-benchmarking-task-verification-in-ocr-reasoning", "title": "Beyond Blind Compliance: Benchmarking Task Verification in OCR Reasoning", "summary": "Researchers introduced VeriOCRBench, a 1,800-sample human-verified benchmark with 1,600 trap-injected invalid tasks across 8 trap types and 200 trap-free controls, to test whether multimodal large language models can detect unanswerable OCR tasks before responding. Evaluating 15 leading MLLMs, the study found persistent blind compliance, diagnosis failures, and prompt-induced over-refusal, exposing a critical reliability gap in current OCR reasoning systems. The code is available on GitHub.", "body_md": "arXiv:2609.00232v1 Announce Type: new\nAbstract: Multimodal Large Language Models (MLLMs) have achieved strong performance on OCR-centric document understanding and text-rich visual reasoning benchmarks. Yet existing evaluations largely assume that every task is valid and answerable. In real-world OCR scenarios, this assumption often fails: questions may rely on illegible text, occluded evidence, nonexistent visual targets, contradictory premises, or missing variables. We study this reliability gap as OCR-grounded Task Verification: before answering, a model should determine whether the Image Premise (IP), Textual Premise (TP), and Question (Q) jointly define an executable task.\nWe introduce VeriOCRBench, a 1,800-sample human-verified benchmark built from source images drawn from 8 OCR-related datasets and spanning 8 real-world image domains, with controlled, image-grounded diagnostic tasks.\nIt contains 1,600 trap-injected invalid tasks across 8 trap types and four verification dimensions---Visual, Contextual, Factual, and Logical---plus 200 trap-free controls for measuring over-refusal. Built with a Visual Atomic Fact (VAF)-anchored pipeline and full human auditing, VeriOCRBench enables decoupled evaluation of task verification, root-cause diagnosis, and over-refusal. Evaluating 15 leading MLLMs reveals persistent blind compliance, diagnosis failures, and prompt-induced over-refusal, exposing a critical reliability gap in current OCR reasoning systems.\nThe code is available at: https://github.com/zy001122/Beyond-Blind-Compliance.", "url": "https://wpnews.pro/news/beyond-blind-compliance-benchmarking-task-verification-in-ocr-reasoning", "canonical_source": "https://arxiv.org/abs/2609.00232", "published_at": "2026-09-02 04:00:00+00:00", "updated_at": "2026-09-02 04:23:00.842880+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-research", "ai-safety"], "entities": ["VeriOCRBench", "Multimodal Large Language Models (MLLMs)", "Visual Atomic Fact (VAF)", "arXiv", "GitHub"], "alternates": {"html": "https://wpnews.pro/news/beyond-blind-compliance-benchmarking-task-verification-in-ocr-reasoning", "markdown": "https://wpnews.pro/news/beyond-blind-compliance-benchmarking-task-verification-in-ocr-reasoning.md", "text": "https://wpnews.pro/news/beyond-blind-compliance-benchmarking-task-verification-in-ocr-reasoning.txt", "jsonld": "https://wpnews.pro/news/beyond-blind-compliance-benchmarking-task-verification-in-ocr-reasoning.jsonld"}}