# AI Firms Buying Old Books: The Hidden Cost of Training Data

> Source: <https://promptcube3.com/en/news/4538/>
> Published: 2026-07-31 12:23:32+00:00

# AI Firms Buying Old Books: The Hidden Cost of Training Data

Let me be clear: I'm not against efficient data collection. Anyone who has built a retrieval pipeline knows the pain of PDFs with garbled text. A physical book scanned with a proper sheet-feed scanner can give you high-quality text that's already segmented. For a company training a model on real-world knowledge, that's gold. If the alternative is crawling sketchy ebook pirate sites, then buying a used bookstore's inventory feels almost ethical.

The "destroy" part is what unsettles me. True, some of these books are in terrible condition — yellowed pages, glue crumbling, covers half-detached. They're not museum pieces. But scanning a book with the intent to discard it means we're permanently losing the chance to preserve the physical object. Libraries and archives care about multiple copies: they add marginal notes, bindings, ownership marks, evidence of how people read.

[Next LinkedIn's new "Seems like AI slop" button: too little, too late? →](/en/news/4536/)

## All Replies （0）

No replies yet — be the first!
