{"slug": "train-as-we-say-not-as-we-do", "title": "Train as We Say, Not as We Do", "summary": "Sony Music Publishing and Warner Chappell Music have filed a copyright lawsuit against Anthropic and its co-founders, Dario Amodei and Benjamin Mann, alleging the company illegally acquired thousands of copyrighted songs through torrents, web scraping, and third-party datasets to train its Claude model, with potential statutory damages reaching billions of dollars. Anthropic disputes the claims. The case follows Anthropic's $1.5 billion settlement in Bartz v. Anthropic over book copyrights, and highlights the legal distinction between how training data is acquired versus how it is used.", "body_md": "TL;DR — Key Takeaways\n\n- Sony Music Publishing and Warner Chappell accuse Anthropic of obtaining thousands of copyrighted songs through torrents, scraping and other sources to train Claude, potentially exposing the company to billions in damages. Anthropic disputes the allegations.\n- The lawsuit highlights an increasingly important distinction between\n**how copyrighted material is acquired, how it is used for AI training and what a model ultimately outputs**—three issues that may require different legal tests. - Anthropic’s criticism of Chinese AI labs for allegedly distilling Claude creates a wider industry dilemma: AI companies want broad freedom to learn from human-created material while demanding strong protection when competitors learn from their models.\n\nClaude may know the words, but Sony and Warner want to know where it learned them—and whether Anthropic paid for the lesson.\n\nSony Music Publishing, Warner Chappell Music and affiliated publishers have filed a sprawling copyright lawsuit against Anthropic and its co-founders, Dario Amodei and Benjamin Mann. The [publishers’ complaint](https://www.musicbusinessworldwide.com/files/2026/08/COMPLAINT-in-Sony_Music_Publishing_US_LLC_e.pdf) accuses them of illegally acquiring thousands of copyrighted musical compositions through torrents, web scraping, third-party datasets and other sources to help build Claude.\n\nThe songs named include “Ain’t No Mountain High Enough,” “All I Want for Christmas Is You,” “Eye of the Tiger,” “September” and Taylor Swift’s “Paper Rings.” The publishers are seeking statutory damages of as much as $150,000 per infringed work, plus potential damages for allegedly removing or altering copyright-management information. With tens of thousands of compositions potentially involved, the theoretical exposure could reach billions of dollars.\n\nThese are allegations, not findings. Anthropic disputes the claims and says it will defend itself.\n\nAlthough the complaint alleges that Claude can produce verbatim or near-verbatim lyrics, that is not the entire case or even its primary foundation. The case is mainly about the provenance of Claude’s training data: where Anthropic obtained copyrighted material, how it copied that material into its library and training datasets, and whether it removed identifying information along the way.\n\nClaude’s ability to reproduce lyrics supports the publishers’ claim that their works entered the training data. It may also provide a separate output-infringement theory. But the publishers are not merely suing because Claude can finish the chorus. They are suing over how the song allegedly got into Claude in the first place.\n\n### The $1.5 Billion Settlement Didn’t End the Story\n\nAnthropic has been here before, although the legal terrain was somewhat different.\n\nIn Bartz v. Anthropic, authors sued over millions of books Anthropic used in developing Claude. Judge William Alsup’s [fair-use ruling](https://docs.justia.com/cases/federal/district-courts/california/candce/3%3A2024cv05417/434709/231) separated the use of the books for training from the means Anthropic used to acquire them.\n\nAlsup found that using copyrighted books to train a model capable of generating new material could qualify as transformative fair use. He also concluded that Anthropic could purchase physical books, remove their bindings, scan them and retain digital copies. Converting a lawfully purchased print library into a digital one was a fair use in the circumstances before him.\n\nWhat Anthropic could not do was make its initial theft disappear by pointing to a transformative use farther down the line. The company had downloaded millions of books from Library Genesis and Pirate Library Mirror, retained them in a permanent central library and paid the copyright holders nothing. A future training use did not retroactively sanitize the acquisition of the source copies.\n\nAnthropic ultimately agreed to a $1.5 billion settlement with authors and publishers whose books were included in the pirated collections.\n\nBut a digital book can contain more than one separately copyrighted work. It may include photographs, illustrations, essays, lyrics or sheet music controlled by rights holders other than the owner of the book copyright. Paying the author or publisher of a songbook does not necessarily settle the claims of the company that owns the compositions printed inside it.\n\nSony and Warner are returning to the same alleged library but asserting rights in another layer of its contents. The $1.5 billion settlement may not have closed Anthropic’s piracy problem. It may have produced a map that other copyright owners can follow.\n\n### What If Anthropic Bought the Music?\n\nThe lawsuit nevertheless contains a difficult question for the music publishers.\n\nSuppose Anthropic legally purchased a physical songbook, destroyed it during scanning and trained Claude on the resulting digital copy. Why would that be legally different from the purchased books covered by Alsup’s ruling?\n\nAnthropic has a credible fair-use argument in that situation. Copyright gives a rights holder control over protected expression, but it does not necessarily give that owner an absolute right to prevent anyone from learning from a lawfully obtained copy. Alsup compared model training to a reader learning from books and then producing something different.\n\nTorrenting an unauthorized copy from a pirate library is much easier to classify. Scraping lyrics from an authorized website is more complicated. The website may possess a license to display the lyrics, but that does not mean every visitor receives the right to copy, retain and reuse them for any commercial purpose. It also does not automatically mean that an internal training use cannot qualify as fair use.\n\nAcquisition, training and output need to be examined separately. Purchasing a songbook is not the same act as scraping a licensed service, and neither is equivalent to downloading a pirate copy. A transformative training use should not retroactively legalize theft. At the same time, ownership of a copyright should not become an automatic veto over every machine that learns from a lawfully obtained work.\n\nOutput raises another independent question. Even if training on a legally acquired song is fair use, that conclusion does not necessarily authorize Claude to distribute the lyrics as a substitute for licensed services. Fair-use training should not become a blanket defense for everything the resulting model produces.\n\n### The Other Models in the Room\n\nWhy, then, do the music-publisher cases appear to be concentrating on Anthropic?\n\nThe answer may not be that Anthropic alone used questionable data. It may be that litigation has exposed unusually detailed evidence about Anthropic’s conduct. The publishers have names, dates, internal communications, source libraries and descriptions of the company’s collection process. They are not filing a complaint based solely on Claude knowing a few lines from a popular song.\n\nOther frontier laboratories face related questions.\n\n[A public court order](https://law.justia.com/cases/federal/district-courts/new-york/nysdce/1%3A2023cv08292/606655/1018/) says an OpenAI employee downloaded pirated books from LibGen in 2018. Those books were used to create the Books1 and Books2 datasets that helped train GPT-3 and GPT-3.5. OpenAI says it later deleted the datasets.\n\nPublishers have accused Google of training Gemini with material obtained from pirated sources and with books Google originally received through Google Books and Google Play for more limited purposes. Meta faces litigation and a growing documentary record involving torrented books allegedly used to train Llama. Authors have also sued xAI, claiming Grok’s broadly described internet training corpus included unauthorized and pirated works.\n\nThese allegations are not interchangeable, and they have not all been proven. The datasets, models, acquisition methods and possible defenses differ. But Anthropic should not be mistaken for the only laboratory facing training-data questions simply because it currently has the most visible stack of evidence against it.\n\nAnthropic may not be the only lab that crossed the line. It may be the first one forced to show exactly where the line was crossed.\n\n### Train as We Say, Not as We Do\n\nAll of this would be complicated enough without Anthropic simultaneously accusing Chinese AI laboratories of doing to Claude something remarkably similar to what publishers say Anthropic did to them.\n\nIn February, Anthropic said DeepSeek, Moonshot AI and MiniMax created approximately 24,000 fraudulent accounts and generated more than 16 million exchanges with Claude. According to [Anthropic’s account](https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks), the Chinese laboratories collected Claude’s responses to distill its reasoning, coding, tool-use and agentic capabilities into their own models.\n\nDistillation is a standard AI-development technique. A more capable model generates outputs that become training material for another model, often allowing the second model to acquire useful capabilities at far lower cost. Anthropic uses distillation within its own development work and acknowledges that the practice can be legitimate.\n\nIts objection is to competitors allegedly accessing Claude through fraudulent accounts and proxy services, concealing their identities, evading regional restrictions and using the results to build competing general-purpose models. Anthropic calls this illicit, industrial-scale extraction. It argues that distillation allows Chinese developers to narrow the capability gap without making the same investments in research, infrastructure and compute.\n\nAnthropic wants the United States to crack down on the practice. It particularly objects to Chinese laboratories using American frontier models to overcome some of the limitations created by restrictions on advanced chips.\n\nThe complaint sounds familiar because it is familiar.\n\n### Training for Me, Theft for Thee\n\nSony and Warner allege that Anthropic acquired valuable human-created works without permission and used them to help build Claude. Anthropic alleges that Chinese laboratories acquired valuable Claude outputs without permission and used them to build competing models.\n\nThe publishers say Anthropic avoided licensing costs. Anthropic says the Chinese labs avoided much of the expense required to develop frontier capabilities independently. Both sides describe industrial-scale operations designed to extract someone else’s investment and convert it into a commercial advantage.\n\nThere are real legal differences. Sony and Warner own registered copyrights in musical compositions. Anthropic’s claims against the Chinese laboratories rely heavily on alleged fraud, circumvention, contractual restrictions, competitive extraction and national-security concerns. Claude’s capabilities are not legally equivalent to the lyrics of “Eye of the Tiger.”\n\nBut the philosophical contradiction remains.\n\nWhen copyrighted books and songs become Claude training data, Anthropic emphasizes transformation and learning. When Claude’s responses become training data for another model, Anthropic emphasizes extraction and theft. Apparently, whether training is innovation or appropriation depends partly on whose model is doing the training.\n\nAnthropic would point to the Chinese labs’ alleged use of fraudulent accounts and deliberate evasion of access restrictions. That is a legitimate argument. A company should be able to enforce reasonable conditions on access to its commercial services.\n\nIt is also an uncomfortable argument for a company accused of torrenting millions of files from known pirate libraries, scraping sites that paid to license their content and stripping identifying information from collected material.\n\nThe point is not that two wrongs make a right. Nor is it that Chinese laboratories should have unrestricted permission to mine Claude. Anthropic invested enormous amounts of money, talent, compute and time in developing its models. It has a legitimate interest in protecting that investment.\n\nSo do the authors, songwriters and publishers whose work helped make these models possible.\n\nThe ownership system frontier AI companies appear to prefer is permissive below their models and protective above them. Human-created material should be widely available for ingestion, transformation and commercialization. Model-created material, however, should be guarded against extraction by anyone hoping to build a competitor.\n\nThat is not a sustainable principle. It is a business preference masquerading as an intellectual-property framework.\n\n### The Rule AI Actually Needs\n\nA workable system has to begin by separating acquisition from training. A transformative purpose may justify certain uses of a lawfully acquired work. It should not retroactively legalize theft used to obtain the source material.\n\nThe system must also distinguish purchasing a copy, accessing a licensed service and torrenting an unauthorized file. All three may place information in front of a person or machine, but they do not create the same rights or obligations.\n\nOutputs should be evaluated independently. Fair-use training should not excuse a model that distributes meaningful substitutes for copyrighted works. At the same time, the occasional presence of protected expression in an output should not automatically make the entire training process unlawful.\n\nSimilar principles should apply to model distillation. Fraudulent accounts, concealed identities and deliberate circumvention can be prohibited without declaring that every act of learning from another model is theft. If transformative learning is a legitimate principle, frontier laboratories cannot claim it exclusively for themselves.\n\nFinally, AI developers should be required to maintain meaningful records showing where their training material originated. An industry building products worth hundreds of billions of dollars should not be permitted to treat the provenance of its essential raw material as unknowable or irrelevant.\n\nThe same basic rules should apply whether the party doing the training is Anthropic, OpenAI, Google, Meta, xAI or a Chinese open-weight developer. They should also recognize the legitimate interests of the authors, musicians, photographers and publishers whose work provided the foundation.\n\nThe AI industry cannot sustain a system in which everything below a frontier model is available for ingestion while everything above that model is protected from extraction.\n\nAnthropic is entitled to protect what it built. But before it tells the rest of the world how models may be trained, it may have to answer more fully for how its own model learned the words.", "url": "https://wpnews.pro/news/train-as-we-say-not-as-we-do", "canonical_source": "https://techstrong.ai/features/train-as-we-say-not-as-we-do/", "published_at": "2026-09-01 13:10:48+00:00", "updated_at": "2026-09-01 13:25:57.853404+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-policy", "ai-ethics"], "entities": ["Sony Music Publishing", "Warner Chappell Music", "Anthropic", "Dario Amodei", "Benjamin Mann", "Claude", "Bartz v. Anthropic", "William Alsup"], "alternates": {"html": "https://wpnews.pro/news/train-as-we-say-not-as-we-do", "markdown": "https://wpnews.pro/news/train-as-we-say-not-as-we-do.md", "text": "https://wpnews.pro/news/train-as-we-say-not-as-we-do.txt", "jsonld": "https://wpnews.pro/news/train-as-we-say-not-as-we-do.jsonld"}}