AI Firms Buying Old Books: The Hidden Cost of Training Data AI firms are buying old books in bulk to scan for training data, often discarding the physical copies afterward, a practice that raises preservation concerns. The author, writing on a tech news site, notes that while efficient data collection is valuable, destroying books eliminates the chance to preserve marginal notes, bindings, and other historical evidence. The piece highlights the trade-off between practical data needs and cultural preservation. AI Firms Buying Old Books: The Hidden Cost of Training Data Let me be clear: I'm not against efficient data collection. Anyone who has built a retrieval pipeline knows the pain of PDFs with garbled text. A physical book scanned with a proper sheet-feed scanner can give you high-quality text that's already segmented. For a company training a model on real-world knowledge, that's gold. If the alternative is crawling sketchy ebook pirate sites, then buying a used bookstore's inventory feels almost ethical. The "destroy" part is what unsettles me. True, some of these books are in terrible condition — yellowed pages, glue crumbling, covers half-detached. They're not museum pieces. But scanning a book with the intent to discard it means we're permanently losing the chance to preserve the physical object. Libraries and archives care about multiple copies: they add marginal notes, bindings, ownership marks, evidence of how people read. Next LinkedIn's new "Seems like AI slop" button: too little, too late? → /en/news/4536/ All Replies (0) No replies yet — be the first