{"slug": "why-is-anthropic-destroying-books", "title": "Why is Anthropic destroying books?", "summary": "Anthropic PBC, the AI company behind the Claude large language model, engaged in 'Project Panama,' a destructive scanning effort that destroyed physical books to train its AI, according to court documents in Bartz v Anthropic PBC. The company chose this method over securing copyright permissions, leading to a $1.5 billion settlement with authors. The court ruled that using copyrighted material to train an LLM did not constitute infringement, treating it as fair use.", "body_md": "Should we destroy all the books in the world?\n\nAn answer to this question can be found in the court documents of Bartz v Anthropic PBC. The northern California district court case, decided in late July this year, highlighted the improbably named “Project Panama”, one of the AI company Anthropic’s efforts to improve its large language model Claude. “What is Project Panama?” [court exhibit 21](https://www.courtlistener.com/docket/69058235/554/21/bartz-v-anthropic-pbc/) asks, in an internal memo. The answer: “Project Panama is our effort to destructively scan all the books in the world.” The memo advises discretion: “Why use a codename? … [B]ecause we don’t want it to be known that we are working on this.”\n\nDestructive scanning was Anthropic’s solution to a problem: in order to “train” Claude, Anthropic had to procure a large, high-quality language dataset, [preferably one created before 2022](https://futurism.com/artificial-intelligence/ai-companies-destroying-rare-books) and the corrupting influence of generative AI on contemporary text. Claude needed as many language combinations as possible, to improve its ability to predict language outcomes. Anthropic needed data, lots of it, of very high quality. Books, as it happens, remain one of the best sources for complex, high-quality, long-form text. As the [court’s decision](https://copyrightalliance.org/wp-content/uploads/2025/06/Bartz-v.-Anthropic-Order.pdf) relates, Anthropic hoped that books’ “well-curated facts, well-organized analyses, and captivating fictional narratives” would help “Claude write as accurately and as compellingly as Authors”.\n\nAnthropic had a choice: it could have secured copyright permission to use existing e-books. This would have required the “legal/practice/business slog”, as Anthropic’s co-founder and CEO [phrased it](https://www.nytimes.com/2025/09/05/technology/anthropic-settlement-copyright-ai.html), of managing copyright. Rather than engage with the texts’ owners, Anthropic first chose to use pirated sources instead, a decision informing the company’s [$1.5bn out-of-court settlement](https://www.washingtonpost.com/technology/2025/09/05/anthropic-book-authors-copyright-settlement/) with authors. When that approach seemed too complicated (or, as the court decision phrases: “Anthropic became ‘not so gung ho about’ training on pirated books ‘for legal reasons’”), Anthropic turned to destructive scanning. As it happened, the judge ruled that using proprietary material to “train” an LLM did not, in and of itself, constitute an infringement of copyright. To the court, it would seem, “training” a corporate product is equivalent to training any human, teaching how to read in order to learn how to write.\n\nShould we be surprised that destroying printed texts seemed easier to Anthropic than working with their human authors? Either way, the decision to use destructive scanning turned the issue into one of logistics. Under US copyright law, the “fair use” doctrine allows you to make “transformative” use of copyrighted works without the owner’s permission. Anthropic took printed books and scanned them, “transforming” or remediating them into a new, electronic format. They then disposed of the original printed copy: the “destructive” part of destructive scanning. Along the way, Anthropic’s vendors had already [sliced the spines and edges of the books](https://www.404media.co/ai-companies-are-buying-tons-of-old-books-because-theyre-free-of-ai-slop/), to scan them more easily before destroying them. “One replaced the other,” as Judge William Alsup wrote, [noting](https://deadline.com/wp-content/uploads/2025/06/anthropic.pdf): “There is no evidence that the new, digital copy was shown, shared, or sold outside the company.”\n\nAnthropic hired an experienced logistics manager, sourced the books from vendors, hired staff and housed the books in a warehouse, and found a digitization vendor to take on the project of the destructive scanning itself. The case exhibits show warehouses of books neatly stacked and labelled on shelves, staff moving between them. As the court’s decision [relates](https://deadline.com/wp-content/uploads/2025/06/anthropic.pdf), Anthropic’s vendors “stripped the books from their bindings, cut their pages to size, and scanned the books into digital form – discarding the paper originals”. In the court’s images, stacks of books await destructive scanning. Not seen: the clean-up project of shredding and disposing of the books (or, “all the books in the world”, reformatted as recycling or landfill).\n\n[Bookishness](https://cup.columbia.edu/book/bookishness/9780231195133/), or the range of meanings attached to [the book as cultural object](https://www.hachettebookgroup.com/titles/leah-price/what-we-talk-about-when-we-talk-about-books/9781541673908/?lens=basic-books), has always lived alongside the book’s role as textual instrument. Witness the [example of images](https://www.nytimes.com/2020/09/18/technology/no-trump-did-not-hold-the-bible-upside-down-at-lafayette-square.html) of Donald Trump, holding up a copy of the Bible at St John’s during the protests of 1 June 2020. Yet unlike many other countries, the US has very few legal provisions relating to the regulation or export of American cultural heritage, and few to none governing the treatment of books.\n\nIn the context of the court’s decision in Bartz v [Anthropic](https://www.theguardian.com/technology/anthropic) PBC, destructive scanning is an effective mechanism to strip authorial involvement from printed texts, in order to use the content of those works to improve the functions of an LLM. It is both legal and less regulated than strip mining. What are the consequences if, as seems likely, this practice is adopted by other generative AI companies, now and in the future? How many warehouses of destructively scanned books would be too many? There is no endangered list for printed works, and little regulation of what might constitute survival of the rare or unique. Still further, there is no formal understanding of the human-generated textual object, in and of itself, as a category of cultural asset or heritage that might require protection. If we take seriously the 2022 threshold, as a moment when AI-generated text began to make significant entry into the textual record, should we start to think of the “wholly human author” as an emergent category of collections preservation and stewardship?\n\nThere is more to say here (what to make, for instance, of the “forever” research library Anthropic states it intends to create), but let me close by observing that Bartz v Anthropic PBC is also a powerful statement on the importance of books or long-form text – and of readers. Judge William Alsup noted: “For centuries, we have read and re-read books. We have admired, memorized, and internalized their sweeping themes, their substantive points, and their stylistic solutions to recurring writing problems.” We should worry that Anthropic decided it was easier to scan and destroy physical books than to deal with the “legal/practice/business slog”. We should worry that the current understanding of fair use allowed Anthropic to decide that it was easier to buy and destroy “all the books in the world” than to pay the creators of those works.\n\nBut perhaps the most telling aspect of this case is that it was so important to Anthropic to have access to a dataset of complex, long-form, uncorrupted text. Let’s ask ourselves why that text should seem so profitable. One response to Bartz v Anthropic PBC might be to refuse to devalue our lives as readers and writers, to claim ownership of the cultural spaces in which our thought and words are created, shared and preserved. The risk with generative AI, as this single instance with Anthropic indicates, is that we cede the means of production of our large language lives: that we turn from creators to consumers, and hand the generative promise of our work to large language models and their proprietors.\n\n-\nKathryn James is the rare book librarian at the Lillian Goldman Law Library at Yale University", "url": "https://wpnews.pro/news/why-is-anthropic-destroying-books", "canonical_source": "https://www.theguardian.com/commentisfree/2026/aug/05/anthropic-ai-destroying-books", "published_at": "2026-08-05 12:05:22+00:00", "updated_at": "2026-08-05 12:23:20.363828+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-ethics", "ai-policy"], "entities": ["Anthropic PBC", "Claude", "Project Panama", "Bartz v Anthropic PBC", "William Alsup"], "alternates": {"html": "https://wpnews.pro/news/why-is-anthropic-destroying-books", "markdown": "https://wpnews.pro/news/why-is-anthropic-destroying-books.md", "text": "https://wpnews.pro/news/why-is-anthropic-destroying-books.txt", "jsonld": "https://wpnews.pro/news/why-is-anthropic-destroying-books.jsonld"}}