{"slug": "old-ocr-text-cripples-language-model-training-and-finebooks-wants-to-fix-that-at", "title": "Old OCR text cripples language model training, and FineBooks wants to fix that at scale", "summary": "Hugging Face and EleutherAI's FineBooks project tested 14 open-source OCR models on more than 2,000 historical book pages, finding the top model, dots.mocr, achieves 97.6 percent character accuracy at under two dollars per thousand pages. The team says this is sufficient for AI training data but not yet for scholarly transcriptions.", "body_md": "The FineBooks project from Hugging Face and EleutherAI tested 14 open-source OCR models on more than 2,000 historical book pages. The top model, dots.mocr, hits 97.6 percent character accuracy at under two dollars per thousand pages. That's good enough for AI training data, but not yet for scholarly transcriptions, the team says.\n\nThe article [Old OCR text cripples language model training, and FineBooks wants to fix that at scale](https://the-decoder.com/old-ocr-text-cripples-language-model-training-and-finebooks-wants-to-fix-that-at-scale/) appeared first on [The Decoder](https://the-decoder.com).", "url": "https://wpnews.pro/news/old-ocr-text-cripples-language-model-training-and-finebooks-wants-to-fix-that-at", "canonical_source": "https://the-decoder.com/old-ocr-text-cripples-language-model-training-and-finebooks-wants-to-fix-that-at-scale/", "published_at": "2026-08-10 18:20:39+00:00", "updated_at": "2026-08-10 18:21:57.948701+00:00", "lang": "en", "topics": ["machine-learning", "ai-research", "ai-infrastructure"], "entities": ["Hugging Face", "EleutherAI", "FineBooks", "dots.mocr"], "alternates": {"html": "https://wpnews.pro/news/old-ocr-text-cripples-language-model-training-and-finebooks-wants-to-fix-that-at", "markdown": "https://wpnews.pro/news/old-ocr-text-cripples-language-model-training-and-finebooks-wants-to-fix-that-at.md", "text": "https://wpnews.pro/news/old-ocr-text-cripples-language-model-training-and-finebooks-wants-to-fix-that-at.txt", "jsonld": "https://wpnews.pro/news/old-ocr-text-cripples-language-model-training-and-finebooks-wants-to-fix-that-at.jsonld"}}