cd /news/machine-learning/old-ocr-text-cripples-language-model… · home topics machine-learning article
[ARTICLE · art-90850] src=the-decoder.com ↗ pub= topic=machine-learning verified=true sentiment=· neutral

Old OCR text cripples language model training, and FineBooks wants to fix that at scale

Hugging Face and EleutherAI's FineBooks project tested 14 open-source OCR models on more than 2,000 historical book pages, finding the top model, dots.mocr, achieves 97.6 percent character accuracy at under two dollars per thousand pages. The team says this is sufficient for AI training data but not yet for scholarly transcriptions.

read1 min views1 publishedAug 10, 2026

The FineBooks project from Hugging Face and EleutherAI tested 14 open-source OCR models on more than 2,000 historical book pages. The top model, dots.mocr, hits 97.6 percent character accuracy at under two dollars per thousand pages. That's good enough for AI training data, but not yet for scholarly transcriptions, the team says.

The article Old OCR text cripples language model training, and FineBooks wants to fix that at scale appeared first on The Decoder.

── more in #machine-learning 4 stories · sorted by recency
── more on @hugging face 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/old-ocr-text-cripple…] indexed:0 read:1min 2026-08-10 ·