{"slug": "dococr-eval-a-correction-based-framework-for-ocr-tool-selection-without-ground", "title": "DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth", "summary": "Researchers propose DocOCR-Eval, an annotation-free framework for automatically evaluating and selecting OCR tools without ground-truth labels. The framework uses a three-stage correction and ranking strategy that aggregates outputs from multiple multimodal large language models to approximate annotation-based tool ordering. Experiments show reliable OCR tool selection can be achieved in label-limited settings across diverse document collections.", "body_md": "arXiv:2607.16203v1 Announce Type: new\nAbstract: Document parsing is a foundational step for document understanding tasks such as visual question answering and key information extraction, as it transforms unstructured scanned images into structured representations by extracting textual, visual, and layout information. While numerous Optical Character Recognition (OCR) engines and multimodal large language models (MLLMs) have been developed for this purpose, selecting an appropriate document parsing solution for a given document collection remains challenging, particularly in label-scarce settings. In this work, we conduct a systematic evaluation of text recognition performance across a diverse set of OCR engines and state-of-the-art MLLMs on multiple scanned document benchmarks spanning different domains and languages. Motivated by the limited contextual reasoning capabilities of many OCR engines and the high cost of manual annotations, we propose DocOCR-Eval, an annotation-free evaluation framework for automatic OCR assessment and selection. DocOCR-Eval employs a three-staged correction and ranking strategy to approximate annotation-based tool ordering without ground-truth labels. We show that aggregating across multiple MLLMs progressively improves alignment with annotation-based rankings. Extensive experiments further demonstrate that reliable OCR tool selection can be achieved in realistic, label-limited settings, providing practical guidance for deploying document parsing systems across diverse real-world document collections.", "url": "https://wpnews.pro/news/dococr-eval-a-correction-based-framework-for-ocr-tool-selection-without-ground", "canonical_source": "https://arxiv.org/abs/2607.16203", "published_at": "2026-07-21 04:00:00+00:00", "updated_at": "2026-07-21 04:11:44.364307+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "computer-vision", "natural-language-processing"], "entities": ["DocOCR-Eval", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/dococr-eval-a-correction-based-framework-for-ocr-tool-selection-without-ground", "markdown": "https://wpnews.pro/news/dococr-eval-a-correction-based-framework-for-ocr-tool-selection-without-ground.md", "text": "https://wpnews.pro/news/dococr-eval-a-correction-based-framework-for-ocr-tool-selection-without-ground.txt", "jsonld": "https://wpnews.pro/news/dococr-eval-a-correction-based-framework-for-ocr-tool-selection-without-ground.jsonld"}}