I was horrified but unsurprised to read reports 1 this week of the AI/LLM
companies buying, butchering, scanning, and shredding rare books at an unprecedented scale.
2With all online forums, PDFs, articles, and books consumed, the next best source of human writing is the abundance of physical books published before the advent of LLMs. The year 2021 3 and the emergence of LLM
text generation is coarsely termed
2“the ensh**tification of all knowledge”. Writing published after this point is potentially LLM-written and therefore unsuitable to be used as training data. Whatever we wrote before this historic point is exclusively the work of human hands and minds.
The race is now on to consume and destroy (at least one copy of) every physical book on the planet first - and I wouldn’t rule out a scenario similar to the global memory shortage of 2025. At this time, purchases became weapons used to strangle the supply chains of competitors and gain an exclusive material advantage. Memory prices shot far out of reach for normal people as AI/LLM companies rushed to secure the remaining global supply of DRAM chips. Will we see the same thing happen with rare, old, and out-of-print books?
“Many companies, including Anthropic, have turned to ingesting physical books instead, which they can buy countless used copies of on the cheap. According to the settled lawsuit, Anthropic used a hydraulic powered cutting machine to neatly remove the pages from the books it procured from book resellers and then scanned them using industrial-grade imaging equipment. In other words, it was literally ripping off authors’ books to train its AI.”
–
Frank L.Futurism[1] The problem here is not just the destruction of rare and priceless books which, though currently unused, are presently accessible human knowledge - hoarding this rare knowledge behind walled gardens is the real issue. Countless dystopian sci-fi visions of humanity’s darkest paths describe the gathering and destruction of knowledge so the record can be set straight by a central authority. Ray Bradbury’s Fahrenheit 451 and George Orwell’s Nineteen Eighty-Four both come to mind as keen examples of governments and their proxies working tirelessly to destroy the past and rewrite history.
Ultimately, this leads to a nightmare situation where all knowledge is gathered by the tech companies feeding these insatiable large language models. Every copy of every book will be frantically secured to keep them out of the hands of competitors. Knowledge will then become accessible exclusively through the censored and monitored outputs from these models. Information deemed ‘unsafe’ will be hidden forever underneath the safety, compliance, and ethics layers of the largest models - governing and correcting the output of every individual response. Whatever is deemed harmful by regulators will be unavailable for the greater good. History will be gatekept.
An argument could be made that these books are much better put to use in this manner. By digitizing and using the books as training data, the knowledge is made more accessible to the everyman who uses ChatGPT as their primary knowledge source. I would be more excited about this if the scans and content were not to be permanently sequestered in each LLM company’s knowledge silo. No guarantee is provided that any of this material will be presented in an impartial manner, or ever released to the public - this would remove the expensive knowledge advantage the particular LLM company gained through the acquisition and scanning process.
“One small book seller said that in April, he suddenly went from selling no more than 20 books a week to
hundreds, and he’s almost certain that the customers are AI labs, noting the random selection of the books and how theyall have ISBNs. He added that his inventory is full with rare and out-of-print books, meaning that an AI company could be destroying some of the few remaining copies that can be found.”–
Frank L.- Futurism[1] “Obviously, printed books represent a treasure trove of information for AI. However, many debate the ethics of removing books from circulation since it is uncertain whether AI companies filter the rare or even out-of-print books from the common titles during digitalization. The other major issue is that scanned books go directly into a private database to train AI, which the general public does not have access to. True, we will have smarter AI, but at the cost of the information not being available to future generations.”
–
Zhiye L.- Tom’s Hardware[4] What can you do about it? Learn to read, then buy and hold paper books. Your children will thank you, and paper books are entirely yours to mark up, borrow, trade, and love - entirely out of the subscription prison most companies are hurriedly constructing today. Paper books can’t be revoked, altered, or monetized past the point of acquisition, and therefore are unattractive to these businesses who seek recurring revenue above all other things.
Is the danger real? Not to a dystopian extent, but just as we’ve all seen the deterioration in the quality of web results - allowing the American technological establishment to act as a middleman between us and all knowledge is a dangerous game. Convenience comes at the cost of narrative injection. Summarized responses at the cost of viewing history through Silicon Valley’s frame. 5 I don’t think Meta and Anthropic are actually bent on buying every paper book in existence, but this month is certainly the opening of an interesting
new frontier of knowledge collection, segregation, and protection. Anything we can collectively do to keep old, beautiful, and rare books out of the scanners and shredders of these LLM leviathans (in particular history books and encyclopedias written before 1945), is a great blessing to our future generations.
Here’s to a future with books!
“AI Companies Are Buying Antique Books, Ingesting Their Contents to Train Models, and Then Destroying Them at Incredible Scale”- Frank Landymore, Futurism, Pub. July 25th 2026,futurism.com/artificial-intelligence/ai-companies-destroying-rare-books↩︎↩︎↩︎LLM:Large language model.A generic term for the algorithms that power ChatGPT, Claude, Microsoft Copilot, Grok, Llama, and many othergenerative pre-trained transformerneural network algorithms. These models arenot AI, but they are branded with the term.↩︎↩︎Some say 2022, but my first conversations with OpenAI’s Davinci model were around August 2021.
↩︎“AI companies are reportedly shredding millions of books after using them to train AI models — tech giants outsource to middlemen to secretly buy up books for training material”- Zhiye Liu, Tom’s Hardware, Pub. July 28, 2026,tomshardware.com↩︎Narrative injection andframe in these sentences both refer to the retrieval of textual information colored with the perspectives and worldview of the silicon valley technology companies. Research performed by Cripps (reported by theSociety for Computers and Law:) revealed that LLMs like ChatGPT and Copilot return*“LLMs are Left-Leaning Liberals: The Hidden Political Bias of Large Language Models”*exclusively left-leaning answers when the neutral option is excluded, and it is difficult to find a mainstream model that will even mention a right-of-center opinion.↩︎