Anthropic to pay $1.5B settlement for pirated books used to train Claude A federal judge has approved a $1.5 billion settlement against Anthropic for using pirated books to train its Claude models, establishing a costly precedent for training-data provenance. The penalty, roughly $3,000 per book across approximately 500,000 works, underscores the critical need for strict copyright compliance in AI training pipelines. Hacker News https://apnews.com/article/ai-anthropic-copyright-settlement-claude-books-bartz-74b140444023898aeba8579b6e9f0d63 Anthropic to pay $1.5B settlement for pirated books used to train Claude Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated. A federal judge has approved a $1.5 billion settlement against Anthropic over the use of pirated books to train its Claude models. This landmark penalty establishes a costly precedent for training-data provenance, which will likely drive up model API costs and force production teams to strictly audit their own fine-tuning datasets for copyright compliance. By turning copyright infringement into a massive balance-sheet liability, this ruling cements data-licensing compliance as a critical, non-negotiable engineering constraint for any team customizing or deploying foundation models. Anthropic will pay $1.5 billion—roughly $3,000 per book across ~500,000 pirated works—for training Claude on illegally obtained copyrighted material, now a court-approved precedent rather than a theoretical risk. If your training or fine-tuning pipeline touches shadow-library sources like LibGen, that data provenance is now a concrete, per-work liability, so audit your corpus sourcing and licensing before it becomes a discovery target.