Anthropic has been telling everyone who’d listen how China is illegally distilling its models, but it turns out that its own hands aren’t particularly clean when it comes to copyright protection.
A federal judge in San Francisco has given final approval to Anthropic’s $1.5 billion settlement with a group of authors who accused the company of using pirated versions of their books to train Claude. US District Judge Araceli Martinez-Olguin signed off on the deal on Monday, dismissing objections from authors who felt the payout was too small, and calling those complaints “not grounded in a realistic assessment of the overall risks and rewards of a trial.”
The number itself is not new. Anthropic and the authors first agreed to the settlement back in September last year, and it was even approved once before by then-Judge William Alsup. What changed this week is that the deal has now cleared its final legal hurdle, with a new judge inheriting the case after Alsup’s retirement and ruling that the settlement stands despite months of authors and publishers arguing it undervalued their work.
At $1.5 billion, this remains the largest copyright settlement in US history, and the first major resolution among the dozens of lawsuits that publishers, news organizations, and artists have filed against AI companies over training data. The math works out to roughly $3,000 per book, covering close to 500,000 titles that Anthropic pulled from pirate repositories including Library Genesis and the Pirate Library Mirror.
How Anthropic ended up here
The case traces back to 2024, when authors Andrea Bartz, Charles Graeber, and Kirk Wallace Johnson sued Anthropic, alleging the company built Claude on the back of pirated books without paying or asking anyone. Alsup’s ruling last June actually went in Anthropic’s favor on the central question, finding that training an AI model on copyrighted books counts as fair use, since the process is transformative rather than a straight reproduction of the original work.
Where Anthropic lost is in how it acquired those books in the first place. Alsup found that the company had downloaded and stored more than 7 million pirated books in what internal documents called a “central library,” a repository that was not necessarily tied to any specific training run and existed independent of the fair use question. That distinction, subtle as it sounds, is what exposed Anthropic to potentially hundreds of billions of dollars in statutory damages had the case gone to trial, which is why the company chose to settle instead.
Anthropic’s deputy general counsel Aparna Sridhar framed the outcome as closure rather than concession, noting that the fair use ruling remains intact and that the settlement was about resolving the piracy claim, not the training claim. More than 91% of eligible authors and publishers have already filed to claim their share of the payout, according to the company.
Not everyone is on board, though. A number of authors and publishers opted out of the class entirely and are pursuing Anthropic separately, including a case tied to a New York Times reporter’s suit against multiple AI firms and one filed by the publisher of the Chicken Soup for the Soul book series. Those cases remain active, and Anthropic’s legal bill for training data is unlikely to stop growing at $1.5 billion.
The distillation irony
The timing of this ruling lands awkwardly for Anthropic, given the company spent much of this year positioning itself as the industry’s leading voice against unauthorized model distillation. In February, Anthropic accused three Chinese labs — DeepSeek, Moonshot AI, and MiniMax — of running coordinated campaigns using roughly 24,000 fake accounts and 16 million queries to extract Claude’s outputs and train cheaper, competing models. The company called on the industry and lawmakers to treat this as a serious threat, citing everything from unfair competition to national security risk.
The response to that campaign was not entirely sympathetic. Critics including Elon Musk pointed out the obvious contradiction: a company built in part on unlicensed books objecting to a rival extracting value from its own model without payment. The counterargument from Anthropic’s defenders is that the two situations are not equivalent, since the Chinese labs were bypassing paid API terms while book authors were never offered any deal at all. Whether that distinction holds up matters less than how it looks, and this week’s ruling does not help Anthropic’s case in the court of public opinion.
For a company that has built much of its public messaging around being the safety-conscious, principled alternative to its rivals, a $1.5 billion admission that it warehoused millions of pirated books complicates that pitch. The legal argument that training on copyrighted material is fair use may have survived intact, but the manner in which Anthropic sourced its training data clearly did not.