AI companies are buying old printed books because clean human text has become scarce, and the fight is no longer only about copyright. It is about who gets to turn books into fuel.
The strangest thing about the new AI data hunt is how physical it has become. After years of scraping the web, labs and their suppliers are back in the ordinary world of shelves, ISBNs, used-book warehouses and bulk purchase orders. Paper is back in fashion.
404 Media reported in July 2026 that ISBNdb, a company better known for book metadata, is now offering high-volume book acquisition services to AI companies. Its pitch is blunt: printed books are valuable because they were written before the web filled up with AI-generated text. That is the commodity now. Not paper. Not bindings. Human language that came before the machine started copying itself.
You can see why AI labs want it. Epoch AI estimated in a 2024 paper that the effective stock of quality-adjusted human-generated public text is about 300 trillion tokens, and that language models could fully use that stock between 2026 and 2032 if current trends continue. Nature published research the same year showing that models trained repeatedly on AI-generated output can degrade over generations. Once you accept those two facts, old books stop looking like dusty leftovers from publishing. They look like inventory.
That doesn't make the practice clean. It makes it profitable.
Ninth Circuit Says Perplexity's Comet Can Shop on Amazon Again The Ninth Circuit Court of Appeals overturned a March injunction that had blocked Perplexity's Comet AI agent from shopping on Amazon, ruling that it's the user, not Perplexity, who legally accesses Amazon's platform. It's the first federal appeals ruling on whether AI agents can shop on third-party platforms without the platform's consent, and... - perplexity comet amazon shopping bot - ai agent computer fraud and abuse
The book supply chain has become training infrastructure #
According to 404 Media, ISBNdb has advertised old printed books as a pool of text still free of synthetic contamination and has described the optics problem to clients directly. The phrase matters because it admits the quiet part. A company can buy a truckload of books legally and still look awful if the public understands that those books may be cut apart and scanned - then simply discarded - so a model can write better email drafts, summaries or chatbot answers.
Anthropic already showed how this works in court. In Bartz v. Anthropic, filings described a program in which the company bought print books and digitized them - then discarded the physical copies. U.S. District Judge William Alsup ruled in June 2025 that training Claude on lawfully acquired books could qualify as fair use, and that converting purchased print books into digital files was also fair use when the print copies were thrown away. That was a major win for AI companies.
But it was not a clean win.
Alsup also ruled that Anthropic could still face claims tied to pirated books. The class action brought by authors Andrea Bartz, Charles Graeber and Kirk Wallace Johnson later produced a $1.5 billion settlement. Reuters, TechCrunch and the Los Angeles Times reported that Judge Araceli Martinez-Olguin granted final approval on July 20, 2026, covering roughly half a million works and paying about $3,000 per eligible work. The lesson for AI companies is narrow but useful: buying and scanning books looks far safer in court than down them from pirate libraries.
Copyright is only part of the problem #
Frankly, the legal question is too small for what is happening here. A judge can say format-shifting is fair use, and the rest of us can still see the cultural trade. Books that once moved from reader to reader now move from seller to scanner to disposal bin. The text survives. The object doesn't.
That may sound sentimental until you get to rare, out-of-print or hard-to-find titles. Libraries preserve physical books because commercial value is a bad measure of long-term value. A book can be obscure and still matter. It can be out of print because the market forgot it, not because the work has no use. When AI data buyers start searching for pre-2022 books at scale, the books least protected by demand may become the easiest to consume.
There is also a market signal here for anyone building or funding AI products. The industry spent years talking as if data were nearly infinite. Now firms are paying for licenses and buying print runs - anything to get at the last clean pools of human text. That tells you more than a keynote does. The constraint is real enough that companies are willing to build procurement systems around it.
Amazon Shuts Its AGI Lab and Cuts Jobs to Chase Enterprise AI Instead Amazon has shut down its AGI Lab and cut jobs across the unit, shelving upgrades to Nova Premier, Omni, Reel and Canvas. The company is redirecting its focus, and its $200 billion 2026 capex budget, toward enterprise AI deployment instead of the frontier model race. - Amazon closes AGI Lab shifts focus - enterprise AI strategy over AGI race
Amazon, Google, Meta, OpenAI, Anthropic, you name it, every serious model company faces the same pressure even if each one sources data differently. The web has been scraped hard. Synthetic data helps in some settings, but it doesn't replace the messy, edited, human record that books contain. If you want models that write like people, you still need people somewhere in the stack.
The next fight will not be about whether AI companies want books. They do. It will be about whether buying a book gives a company the right to destroy the object, keep the text and turn that text into a commercial model that can compete with the people who wrote the books in the first place.
Also read: Researchers Show How Microsoft Copilot Can Turn One Login Into $247,500 • Micron Launches a $250 Million Fund to Bankroll the AI Startups Buying Its Chips • Robotera Is Said to Weigh a Hong Kong IPO to Raise Up to $1 Billion