Is It Legal for AI to Train on Your Data? A 2026 legal analysis finds that training AI models on public, copyrightable work is mostly legal in most jurisdictions, with US fair use doctrine favoring the practice. In the Anthropic books litigation, a court called training "spectacularly" transformative, yet Anthropic still agreed to pay roughly US$1.5 billion — about US$3,000 per work — because acquiring the books from pirated shadow libraries was a separate wrong fair use did not excuse. Somewhere in the training data of the model you used this morning, there is a decent chance you are in there. Not you by name, necessarily, but your Reddit comments, the code you pushed to a public repository, the blog you kept in 2019, the photographs you posted before you thought to wonder where they would end up. The models that answer your questions and finish your sentences were built by reading an enormous amount of what people put on the internet, and “what people put on the internet” includes a great deal of what you put on the internet. The natural question — the one that arrives the moment this stops being abstract — is whether any of that was allowed. Did an AI company need your permission to learn from your work? Did it need to pay you? Was it, in a word, legal? The honest answer in 2026 is that it is mostly legal, in most places, but the law is still being written in courtrooms as we go, and “legal” has quietly come apart from “something you agreed to”. This piece is a map of where the line currently sits, from the point of view of the person whose data it is. We should be precise about scope first, because two very different questions get muddled together. One is about public, published, copyrightable work — your writing, art, music and code — and whether copyright law lets a company train on it. The other is about private data you type into a product — your chatbot conversations, your work documents — and whether the product’s terms let it learn from you. This article is mostly about the first. The second is governed less by copyright and more by a settings toggle, and we have written before about how those toggles tend to be set; the short version is that the default is often to train on your work unless you find the switch https://theaidownside.com/posts/atlassian-trains-its-ai-on-your-work-by-default.html . A model is not a filing cabinet with your blog post in a drawer. Training is the process of adjusting billions of numerical weights so the system gets better at predicting the next token, pixel or frame; your work is one of the examples it learns the patterns from. This is why the companies argue that training is not “copying” in the everyday sense — the finished model does not, in the normal case, contain a retrievable copy of your article. It is also why the critics counter that the whole edifice was nonetheless built on copying your article at least once, to feed it in, and that a machine which can sometimes reproduce a work near-verbatim has not fully “forgotten” it. Both things are true at once, and that tension is exactly what the courts are trying to resolve. It matters because the legal treatment turns on framing. If training is fundamentally about learning uncopyrightable patterns, it looks like fair use. If it is fundamentally about ingesting and storing millions of protected works to build a commercial product that competes with them, it looks like infringement at industrial scale. The same activity supports both stories, and which one prevails is being decided piece by piece. In the United States, the relevant test is fair use, and the clearest signal so far came from the litigation against Anthropic over its use of books. The court found the training itself to be transformative — “spectacularly so”, in the judge’s words — on the reasoning that a model learning from books to produce new text is doing something different from the books’ original purpose. That is about as favourable a statement as the AI industry could have hoped for on the core question. And yet Anthropic still agreed to pay roughly US$1.5 billion, in a settlement given approval in 2026, at an estimated US$3,000 per work. The reason is the crucial distinction the case drew: training on the books might have been fair use, but obtaining them by downloading pirated copies from shadow libraries was a separate wrong that fair use did not excuse. The lesson is not “training is illegal”. It is “the training may be defensible; the piracy on the way in is what costs you”. How the data was acquired has become its own front in the war. The law’s current answer to “can they train on my work?” is a lawyerly “mostly yes — so long as they didn’t pirate it on the way in.” The picture is not uniform, which is the honest and slightly uncomfortable part. In the Thomson Reuters case against Ross Intelligence, a court found that copying legal headnotes to build a rival research tool was not fair use — though it stressed that Ross’s product was a search tool, not a generative model, which may limit how far that ruling travels. In the authors’ case against Meta, a different judge found the training defensible but pointedly criticised the plaintiffs for making a “half-hearted” argument on the one factor that may matter most: market harm. And the highest-profile fight, The New York Times against OpenAI, is still grinding on. The through-line is that American courts are converging not on a yes-or-no rule but on a fact-specific question: does the model’s output substitute for, and erode the market for, the thing it learned from? Cross the Atlantic and the legal furniture changes. There is no broad fair-use doctrine; instead there are specific copyright exceptions, and the relevant one is for text and data mining. Under the EU’s copyright directive, mining copyrighted material — including for commercial AI training — is permitted unless the rightsholder has expressly reserved their rights in a machine-readable way. The AI Act layers transparency duties on top, requiring general-purpose model providers to respect those reservations and publish summaries of what they trained on. Read that carefully and the catch appears. The opt-out is real, but it is built for entities that control a website or a dataset and can attach a machine-readable “do not mine” flag: a stock-photo library, a news publisher, a big platform. It is not built for an individual whose photographs live on someone else’s social network and whose comments are scattered across a dozen forums. You cannot realistically reserve rights you do not technically control, which means the European “opt-out” is, for most ordinary people, a right they have no practical way to exercise. The UK, meanwhile, gave the world its first substantial courtroom test. In Getty Images v Stability AI , decided by the High Court in November 2025, Getty ended up abandoning its main training-related claims during the trial and the court rejected the remaining secondary-infringement argument, holding that a model’s weights are not themselves “infringing copies” because they do not store the original images. Getty won only a narrow, historic trademark point about watermarks appearing in outputs. It was widely read as a win for the AI developer — but a narrow, technical one that left the biggest questions, including whether the initial training in the UK infringed, deliberately unanswered. Amid the litigation, the US Copyright Office weighed in with the third part of its report on AI and copyright in May 2025, and it is worth reading because it is more sceptical than the early court wins suggest. Its position, in essence, was that using vast quantities of copyrighted work to build systems that generate content competing with that work — particularly where the material was accessed illegally — goes beyond the established boundaries of fair use. It spent considerable effort dismantling the strongest arguments the industry had been making, and floated the idea of “market dilution”: the harm not of copying one book, but of flooding the market with cheap machine-made substitutes for a whole category of human work. A report is not a statute and not a court ruling, and its release was politically bruising. But it is a signal that the official copyright authority does not regard the fair-use question as comfortably settled in the industry’s favour, and that the market-harm argument the Meta plaintiffs fumbled is the one that could, argued properly, change the outcome. Strip away the case names and the position of an ordinary person is fairly clear, if not especially comforting. If you have posted something publicly, it has very likely already been used in training, and the current weight of the law does not treat that as wrong in itself. The places the law bites are around the edges — pirated acquisition, near-verbatim regurgitation, provable market harm — and those are fights conducted by well-resourced plaintiffs, not by you as an individual whose comment history became one ten-millionth of a dataset. Two further realities sharpen the point. First, the terms of service you accept increasingly grant training rights up front, which is why a story about a platform quietly opting you into AI training by default https://theaidownside.com/posts/twitch-opts-you-into-amazon-ai-training.html is now a genre rather than an aberration. Consent obtained through a pre-ticked box and a settings page three menus deep is consent in the legal sense and nowhere near it in the meaningful one. Second, the tools that would let you claw data back barely exist. We have written about how hard it is to get your data out of an AI tool https://theaidownside.com/posts/can-you-get-your-data-out-of-an-ai-tool.html once it is in, and the same asymmetry applies here: getting your work into a training set takes a crawler a fraction of a second; getting it out is, for practical purposes, impossible once a model is trained. It is worth separating this cleanly from the related question of what you can own at the other end of the pipe — whether the outputs a model produces from all that training can themselves be copyrighted. That is a distinct problem with its own emerging answer, which we covered in our piece on whether you can copyright what AI makes https://theaidownside.com/posts/can-you-copyright-what-ai-makes.html . Training is about the inputs; authorship is about the outputs; the law is unsettled at both ends. None of this leaves you completely without options, but honesty requires being clear about how limited they are: robots.txt disallow for the named AI crawlers, and the emerging opt-out conventions, are the one place an individual’s wishes are actually legible to the systems doing the scraping. They only work for content on infrastructure you own, and only for crawlers that choose to honour them. The law here is not static, and it is not obviously heading in the companies’ favour forever. The market-harm argument is getting sharper, regulators are more sceptical than the first headlines implied, and every settlement redraws the line a little. But the gap the reader should hold onto is the one between legal and consented to . Right now, a great deal of training on your work is probably lawful, and almost none of it was anything you would recognise as a choice you made. Closing that gap — making legality depend on genuine, informed permission rather than on a crawler’s head start — is the actual argument, and it is far from over. Originally published at theaidownside.com https://theaidownside.com/posts/is-it-legal-for-ai-to-train-on-your-data.html — evidence-first reporting on the costs and trade-offs behind AI products.