cd /news/artificial-intelligence/ai-companies-destroy-millions-of-boo… · home topics artificial-intelligence article
[ARTICLE · art-78363] src=insideai.news ↗ pub= topic=artificial-intelligence verified=true sentiment=↓ negative

AI Companies Destroy Millions of Books to Train Models, Investigation Reveals

AI companies are systematically buying millions of secondhand books, scanning them for training data, and destroying the physical copies, according to an investigation by 404 Media. The practice, which accelerated sharply in April 2026, uses intermediaries to source high-quality, pre-2022 content while avoiding public scrutiny and copyright lawsuits. Anthropic set the precedent by using hydraulic-powered cutting machines to remove pages before scanning, with a federal judge ruling the practice as fair use, though the company still paid $1.5 billion for maintaining pirated copies.

read4 min views5 publishedJul 29, 2026
AI Companies Destroy Millions of Books to Train Models, Investigation Reveals
Image: Insideai (auto-discovered)

July 29, 2026, (Inside AI) — AI companies are systematically buying millions of secondhand books, scanning them for training data, and then destroying the physical copies, according to an investigation by 404 Media. The practice, which accelerated sharply in April 2026, uses intermediaries to source high-quality, pre-2022 content while avoiding public scrutiny and copyright lawsuits.

The core driver is the contamination of internet data with “AI slop,” low-quality machine-generated content that degrades model performance when used for training. Books published before the generative AI boom offer a clean, human-authored alternative. ISBNdb, the world’s largest book database, now markets this advantage directly to AI firms, stating that print books from the pre-LLM era are “structurally guaranteed to be free of this contamination.”

ISBNdb’s website is blunt about the reputational risk. It acknowledges, “The optics problem is real. AI company destroys two million books is not a headline that generates sympathy.” The platform facilitates bulk orders ranging from 1,000 copies to one million books in single transactions, while keeping buyer identities confidential. Anonymous booksellers report weekly sales jumping from about 20 books to several hundred.

Anthropic set the destructive precedent, using hydraulic-powered cutting machines to remove pages from books before scanning them with industrial imaging equipment. A federal judge ruled this practice transformed the original texts into new digital forms, qualifying as fair use. However, Anthropic still paid $1.5 billion for maintaining pirated copies.

Booksellers in the Netherlands have reported targeted bulk purchases of scarce and out-of-print titles, some of which exist nowhere else globally. These acquisitions are raising alarm among preservationists, as unique physical copies are being destroyed for digital ingestion. The practice highlights a tension between advancing AI capabilities and preserving cultural heritage.

How Pre-2022 Books Became AI's Clean Data Source #

The shift toward physical books as training data stems from a well-documented problem: model collapse. Research shows that when AI is trained on AI-generated content, output quality degrades over successive generations. A 2023 study in Nature demonstrated that large language models lose coherence when recursively trained on their own outputs. Pre-2022 books are valued because they are “structurally clean” of modern AI contamination, as ISBNdb’s marketing emphasizes.

ISBNdb’s role is pivotal. The platform, which indexes over 30 million ISBNs, now offers a service specifically for AI training data procurement. Its website states that physical books published before the LLM era are “structurally guaranteed to be free of this contamination,” and it explicitly acknowledges the destructive outcome: “The optics problem is real.” This candid admission underscores the ethical tightrope the industry walks.

The April 2026 surge in purchases coincides with growing demand for high-quality, uncontaminated datasets. As AI companies exhaust easily accessible digital text, they are turning to physical media. This trend mirrors earlier data scraping controversies but adds a physical dimension: the literal destruction of books after scanning. The practice raises questions about the long-term availability of rare texts and the environmental cost of mass book disposal.

Anthropic’s legal victory set a critical precedent. The company argued that physically dismantling books and scanning them constitutes transformative use under copyright law. A federal judge agreed, finding that the process created new digital forms distinct from the originals. However, the ruling did not fully resolve the copyright implications, as Anthropic still settled for $1.5 billion over pirated copies. This suggests that while the physical destruction method may be lawful, the underlying data acquisition can still face legal challenges.

Booksellers are caught in a bind. ISBNdb’s confidentiality promises prevent them from confirming whether buyers are AI labs, but the scale and nature of orders often make it obvious. One Dutch bookseller, speaking anonymously, described orders for “scarce and out-of-print titles” that exist nowhere else. The sense of unease is palpable, as these transactions may permanently remove unique cultural artifacts from public access.

The practice also intersects with broader debates about AI and copyright. Unlike digital text, which can be licensed or scraped, physical book destruction leaves no trace of the original medium. This makes it harder to track what data has been ingested and to enforce attribution or compensation. Critics argue that even if fair use applies, the systematic destruction of books undermines the public interest in preserving knowledge.

Looking ahead, the industry faces a reckoning. As AI companies seek ever-larger datasets, the demand for pre-2022 books may intensify, potentially driving up prices and accelerating the loss of rare editions. Some preservationists are calling for regulations requiring digital copies to be deposited in public archives before physical destruction. Without such measures, the “optics problem” ISBNdb warns about could become a lasting cultural scar.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @404 media 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/ai-companies-destroy…] indexed:0 read:4min 2026-07-29 ·