cd /news/artificial-intelligence/anthropic-destroys-physical-books-fo… · home topics artificial-intelligence article
[ARTICLE · art-85842] src=insideai.news ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Anthropic Destroys Physical Books for AI Training, Court Documents Show

Anthropic has been buying millions of physical books, destroying their bindings, and scanning the pages to train its large language models, according to newly unsealed court documents. The practice, known as destructive scanning, is part of a broader industry scramble for clean, human-written text as web data becomes polluted with synthetic content. Booksellers in Australia and Europe have reported unusual bulk orders, and the trend raises unresolved copyright questions, with legal experts noting it may be a middle path between licensing and shadow libraries.

read3 min views6 publishedAug 4, 2026
Anthropic Destroys Physical Books for AI Training, Court Documents Show
Image: Insideai (auto-discovered)

August 4, 2026, (Inside AI) — Anthropic has been systematically buying millions of physical books, destroying their bindings, and scanning the loose pages to train its large language models, newly unsealed court documents confirm. The practice, known as destructive scanning, is part of a broader scramble by AI companies to secure high-quality, human-written text as training data.

Booksellers in Australia and Europe have reported unusual bulk orders from intermediaries, often for rare or out-of-print titles. The Guardian and Fortune documented cases where sellers were unaware the books would be dismantled for AI training. This marks a shift from earlier revelations about Project Panama, Anthropic’s initiative to amass a vast corpus of digitized books, first reported by the Washington Post in January 2026.

The practice is now rippling up the supply chain, with booksellers noticing demand for obscure works. Intermediaries often shield the end buyer, making it difficult to trace how the books will be used. The trend underscores a growing hunger for clean, pre-AI text as web-scraped data becomes increasingly polluted with synthetic content.

Destructive Scanning Fuels Data Hunger #

Large language models learn from vast text corpora, and books offer long-form, edited content across diverse subjects. ISBNdb, a company marketing book-sourcing services to AI firms, argues that pre-generative AI books are especially valuable because they lack AI-generated text. As synthetic content proliferates online, older physical books are prized for their purity.

Destructive scanning involves cutting book spines and feeding loose sheets into high-speed scanners. The process yields flat, high-quality images for optical character recognition but permanently destroys the original copy. This contrasts with non-destructive methods used by preservation projects like the Internet Archive, which photographs books without dismantling them.

Chris Freeland, Director of Library Services at the Internet Archive, told the Indian Express that the organization’s 2021 explanation of its scanning process “continues to be the best description of its digitisation process today.” The Archive has long avoided destructive methods to protect brittle or rare volumes.

Yet for AI companies, speed and image quality often outweigh preservation. Many books lack clean digital editions, and publishers tightly control commercial e-book access. Physical copies remain a straightforward, if controversial, path to obtaining training data.

The legal landscape is unsettled. Arul George Scaria, professor at the National Law School of India University, described the practice as a middle path between licensing and using shadow libraries. “Buying physical copies and getting them digitised for training appears to be the middle path taken by many AI developers now, as many presume that this is fairer to the authors and publishers,” he said.

Under Indian law, fair dealing exceptions in Section 52 of the Copyright Act of 1957 may apply if digitization is solely for training. But Scaria cautioned that copyright alone may not address author anxieties. He called for solutions beyond law, such as mandatory AI-content labeling and funds for affected creators.

The controversy echoes the Google Books litigation of the 2000s, when Google’s mass digitization project sparked years of copyright battles. Jannis Lennartz, Visiting Professor at Humboldt-Universität zu Berlin, argues in a Verfassungsblog essay that today’s efforts differ in purpose: instead of making books searchable, AI firms convert them into machine-readable training data. This shifts the economics of digitization, treating books as raw material for AI rather than objects of preservation.

Meanwhile, the Internet Archive’s non-destructive approach faces its own legal challenges, including a 2023 ruling that its controlled digital lending program violated copyright. The juxtaposition highlights a deepening divide over how society values physical books in the age of AI.

As the scramble for training data intensifies, the fate of millions of books hangs in the balance. The practice raises urgent questions about cultural preservation, fair use, and the hidden costs of building ever-larger models. For now, booksellers remain on the front lines, often unaware that their sales fuel the destruction of the very works they sell.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @anthropic 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/anthropic-destroys-p…] indexed:0 read:3min 2026-08-04 ·