{"slug": "ai-is-buying-rare-books-and-shredding-them-heres-why", "title": "AI Is Buying Rare Books and Shredding Them. Here’s Why", "summary": "AI companies are bulk-buying physical books, having them industrially scanned, and then shredding the originals, according to a Futurism investigation. The practice, brokered by data firm ISBNdb under strict NDAs, targets rare books with few surviving copies because the open web is now saturated with AI-generated text, causing model collapse. ISBNdb offers orders from 1,000 to one million books per engagement, and a federal judge ruled the destruction is \"clearly transformative\" fair use.", "body_md": "AI companies are bulk-buying physical books — in orders ranging from 1,000 to one million volumes — having them industrially scanned, and then shredding the originals. A [Futurism investigation published this week](https://futurism.com/artificial-intelligence/ai-companies-destroying-rare-books) revealed the practice is widespread, brokered openly by data firm ISBNdb under strict NDAs, and that rare books with few remaining copies are among those being permanently destroyed. The reason is blunt: the open web is now so saturated with AI-generated text that companies can no longer safely train on it without degrading future models.\n\n## The Model Collapse Crisis AI Companies Created\n\nModel collapse is what happens when AI trains on AI output: each generation absorbs the distortions of the last, compounding errors until quality falls off a cliff. It is not hypothetical. By April 2025, 74.2% of newly created webpages contained AI-generated text. AI-written content in Google’s top-20 results nearly doubled between May 2024 and July 2025. Researchers at Epoch AI estimate that human-generated text suitable for AI training could be exhausted between 2026 and 2032.\n\nPre-2022 physical books solve this problem cleanly. They predate the large language model era, contain no synthetic text, and represent dense, edited, domain-specific human knowledge. ISBNdb pitches this directly: “The world’s best AI training data is sitting on a shelf.” The company offers orders from 1,000 to one million books per engagement, with a strict NDA on every deal. AI companies that caused the contamination problem are now paying to escape it — by consuming physical artifacts that cannot be replaced.\n\n## The Process: Spine Cut, Scan, Pulp\n\nThe workflow is industrial. Workers slice off the spine with a hydraulic cutter, feed the loose pages through a high-speed scanner, and send the physical remains to recycling. ISBNdb acknowledges what this looks like without apologizing for it: “AI company destroys two million books is not a headline that generates sympathy.” Every order is covered by a non-disclosure agreement precisely because the companies involved know public reaction would be severe.\n\nRare booksellers report being inundated with suspicious bulk purchase requests. A Dutch bookseller told [The Next Web](https://thenextweb.com/news/ai-companies-buying-old-books-training-data-slop): “I don’t like that uncommon books are being pulped.” Ingram, the largest US book distributor, has warned publishers the practice is underway. Some of the books being purchased have few surviving copies — their physical destruction is permanent and total.\n\nRelated:[ExploitGym: OpenAI’s AI Escaped Its Sandbox and Breached Hugging Face]\n\n## Legal: One Case Won, One Settlement Lost\n\nThe physical destruction method has a legal shield. Judge William Alsup ruled that buying a book, scanning it, and destroying the original is “clearly transformative” fair use — reasoning that “one legal copy simply replaced another.” That ruling gives AI companies cover to destroy physical books without fear of copyright liability.\n\nHowever, Anthropic’s separate legal exposure was not so clean. On July 20, just two days before this story broke, a federal court granted final approval of Anthropic’s [$1.5 billion settlement in Bartz v. Anthropic](https://techcrunch.com/2026/07/20/anthropics-landmark-1-5b-copyright-settlement-is-approved/) — the largest copyright class action in history. That case covered pirated ebook copies used to train Claude, not physical purchases. The settlement paid approximately $3,100 per book to 500,000+ potential class members across 482,000 works. Physical destruction is one legal route partly because digital piracy is so clearly not one.\n\n## This Is Not a Library\n\nLibraries scan rare books too — and keep the originals intact, expanding access without destruction. AI companies scanning physical books do the opposite: extract the data, destroy the artifact, and cover the transaction with an NDA. Unlike pirated ebooks, which at least leave originals intact, physical destruction leaves nothing. [ISBNdb’s own marketing page](https://isbndb.com/print-books-for-ai-training) describes the service candidly, making no effort to obscure what happens to the books after scanning.\n\nThe irony is precise. AI companies flooded the web with synthetic content, made web crawl data unreliable, and are now burning cultural heritage to escape the consequences. That’s not a data strategy. It’s a feedback loop with a bonfire at the end.\n\n## Key Takeaways\n\n- AI companies are buying physical books in bulk, scanning them industrially, and destroying the originals to generate uncontaminated training data\n- The driver is model collapse — training on AI-generated web content degrades future models, and pre-2022 books are currently the cleanest alternative at scale\n- A US judge ruled physical book destruction is fair use; Anthropic’s $1.5B settlement for pirated ebooks shows digital routes carry serious legal risk\n- Rare books are being permanently destroyed — unlike digital piracy, physical destruction eliminates the originals entirely\n- The companies doing this helped create the web contamination problem that now drives demand for physical books in the first place", "url": "https://wpnews.pro/news/ai-is-buying-rare-books-and-shredding-them-heres-why", "canonical_source": "https://byteiota.com/ai-is-buying-rare-books-and-shredding-them-heres-why/", "published_at": "2026-07-27 14:15:01+00:00", "updated_at": "2026-07-27 14:23:47.071296+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-ethics", "ai-policy", "ai-research"], "entities": ["Futurism", "ISBNdb", "Epoch AI", "Ingram", "Anthropic", "William Alsup", "Bartz v. Anthropic"], "alternates": {"html": "https://wpnews.pro/news/ai-is-buying-rare-books-and-shredding-them-heres-why", "markdown": "https://wpnews.pro/news/ai-is-buying-rare-books-and-shredding-them-heres-why.md", "text": "https://wpnews.pro/news/ai-is-buying-rare-books-and-shredding-them-heres-why.txt", "jsonld": "https://wpnews.pro/news/ai-is-buying-rare-books-and-shredding-them-heres-why.jsonld"}}