cd /news/artificial-intelligence/heres-a-balm-if-the-idea-of-destroyi… · home topics artificial-intelligence article
[ARTICLE · art-93842] src=arstechnica.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↓ negative

Here’s a balm if the idea of destroying books to train AI breaks your heart

AI companies are destroying old books to train their models, a practice that book lovers find distressing, but Google patented a non-destructive book-scanning technology in 2009 that could avoid such waste, though it is slower and costlier. The Internet Archive, which helps libraries preserve aging collections, has long understood that scanning old texts requires time and attention, making it a human job.

read1 min views1 publishedAug 12, 2026
Here’s a balm if the idea of destroying books to train AI breaks your heart
Image: Arstechnica (auto-discovered)

If you can truly appreciate an old book—and maybe even marvel at how its fragile, yellowing pages contain some of the earliest ways that people tried to make sense of the world around them—then headlines about tech companies destroying books to train AI likely torture a tender part of your soul. It’s indeed depressing to imagine piles of book spines waiting to be fed into wood chippers while torn-out pages are cropped, scanned, and trashed. But that’s the cheapest and easiest way to scan books as fast as possible, and AI companies are in a race to advance their models by training on the kind of engaging, high-quality, long-form texts that can only be found in books. So book lovers fear it’s likely that the practice is happening on a grander scale than is currently being reported and that some physical copies of books will be lost forever.

What makes this destruction extra painful, though, is that it doesn’t have to be this way.

Google patented a non-destructive book-scanning technology in 2009 that AI firms could use to efficiently scan books—if they were willing to slow down and invest in the process. It’s not perfect, however; studies have found that the curve of the page can distort text, and pages can be missed. As Wired reported, glitches can happen when workers move too quickly, including disembodied hands obscuring pages.

Overall, the trade-offs in cost and speed may not appeal to AI firms looking for the cheapest way to scan millions of titles, and Google’s method may not be the best way to handle rare books anyway. The Internet Archive, which helps libraries preserve aging collections, has long understood that scanning old texts takes time and attention to limit handling, and that’s why it’s considered such a human job.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @google 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/heres-a-balm-if-the-…] indexed:0 read:1min 2026-08-12 ·