The DOJ’s AI Fair Use Brief Is Correct, But From A DOJ That Has No Credibility The U.S. Department of Justice filed a brief in the New York Times' copyright case against OpenAI arguing that AI training constitutes fair use, a position that legal analysts say is correct but comes from an administration with diminished credibility. The DOJ contends that a ruling against fair use would harm competition by limiting AI development to large companies and would disproportionately benefit legacy media. The brief also invokes national security concerns, though critics find that argument weak. The DOJ’s AI Fair Use Brief Is Correct, But From A DOJ That Has No Credibility from the worst-doj-you-know-just-made-a-great-point dept I am never asking the monkey’s paw for a federal government that defends fair use again. For decades on Techdirt, I’ve talked about the absolute necessity https://www.techdirt.com/2015/02/23/reminder-fair-use-is-right-not-exception-defense/ of strong fair use, and have been disappointed over and over again at how little the federal government has fought for fair use. For decades, the rare times when the federal government has weighed in on copyright cases, it’s often been in support of copyright maximalism. So, when the government finally makes a strong stand for fair use… it’s this White House? With this DOJ? And, in defense of a giant centralized AI provider? Sigh. So, look, you’re right to be skeptical about the DOJ weighing in regarding the big OpenAI copyright case filed by the NY Times. You’re right to be skeptical about the DOJ’s motives. You’re right to be skeptical about OpenAI’s motives. But the DOJ is correct in its legal analysis https://storage.courtlistener.com/recap/gov.uscourts.nysd.640396/gov.uscourts.nysd.640396.1682.0.pdf . AI training absolutely is fair use, and a ruling the other way would blow a hole in fair use protections that have nothing to do with AI, including everything from search engines to book scanning to reverse engineering to most importantly text and data mining for research. Also, the NY Times’ case against OpenAI is incredibly weak https://www.techdirt.com/2024/03/05/openais-motion-to-dismiss-highlights-just-how-weak-nyts-copyright-case-truly-is/ , and as I’ve discussed involves bizarre theories of copyright that would put all sorts of companies including the NY Times itself at real risk of liability for doing basic reporting https://www.techdirt.com/2023/12/28/the-ny-times-lawsuit-against-openai-would-open-up-the-ny-times-to-all-sorts-of-lawsuits-should-it-win/ . So, having the federal government weigh in and make good points could actually be helpful. However, this is also why it’s so frustrating that the Trump DOJ and Todd Blanche and Pam Bondi before him completely burnt through the presumption of regularity https://www.techdirt.com/2026/07/31/federal-judges-chastise-trumps-justice-department-for-unlawful-unethical-and-unseemly-conduct/ by repeatedly filing bullshit briefs in bullshit cases. Because when they actually file a reasonable thing in an important case, judges are still going to be quite skeptical. But, in this case, the DOJ is correct. We’ve talked about some of this already, regarding the big case against Anthropic where Judge William Alsup found that AI training was easily fair use https://www.techdirt.com/2025/06/26/judge-alsup-training-ai-on-copyrighted-works-fair-use-building-pirate-libraries-not-so-much/ . Much of the coverage of the DOJ’s filing focuses on the “national security” claims which the filing itself spends much of the opening on — the argument that if we don’t let American companies train on everything, China will eat our lunch in AI. That argument is pretty weak, and it’s also unnecessary. The fair use analysis stands on its own without any appeal to beating our adversaries. But much more interesting and correct to me is the argument that if training is not fair use, then only a few giant, wealthy companies can afford to create AI tools, and we’d just be recreating the broken “big tech” structures of the last decade, rather than enabling more decentralized, more widely competitive tools: An erroneous fair use ruling would hamper competition in the market for LLMs, because only the largest technology companies might have the capital necessary to pay licensing fees. And such licensing fees would disproportionately benefit legacy media outlets due to the sheer volume of their written publications. By contrast, if not hindered by a strained understanding of copyright law, LLMs can and should help level the playing field between mainstream and independent publishers, for several reasons. Authors with limited resources can use LLMs to compete by, for example, using an LLM to generate an image to accompany an article—which otherwise might require a photographer or license . And LLMs can direct users to dissenting sources that offer contrary information or perspectives. It is not in the public’s interest for the largest technology companies to have an oligopoly on LLM training due to licensing entry barriers that function primarily as large subsidies for old mainstream media companies. To me this is the whole ballgame, and part of what makes it so frustrating that many people insist training can’t be fair use. They often think the end result is somehow punishing the “big” AI companies, but the reverse is true. A finding against fair use locks in the biggest AI companies, and wipes out everyone else, especially decentralized open weight models that actually empower end-users without enabling tech giants. The DOJ also cites all the right precedents including some that previous administrations were not happy about on a point that gets mangled constantly: fair use isn’t just a defense you raise after infringing. It means there was no infringement in the first place. A copyright owner’s exclusive rights are thus subject to various exceptions and limitations. For example, “copyright assures authors the right to their original expression, but encourages others to build freely upon the ideas and information conveyed by a work.” Feist, 499 U.S. at 349- 50. “This principle, known as the idea/expression or fact/expression dichotomy, applies to all works of authorship.” Id. at 350. Relatedly, and as most relevant here, the “fair use” doctrine provides that certain secondary uses of a copyrighted work are “not an infringement.” 17 U.S.C. § 107. Although fair use originated as “judge-made,” Congress subsequently codified it. Campbell, 510 U.S. at 576. The statute continues a common-law tradition, which recognized that certain amounts and types of copying must occur to promote the purposes of the Intellectual Property Clause. See id. at 575 citing U.S. CONST. art. I, § 8, cl.8 . Fair use is an “equitable rule of reason that permits courts to avoid rigid application of the copyright statute when, on occasion, it would stifle the very creativity which that law is designed to foster.” Google LLC v. Oracle Am., Inc., 593 U.S. 1, 18 2021 . It also notes correctly, though contrary to the belief of many non-copyright lawyers that fair use was written to be broad and flexible, not just based on a specific set of categories, or if you meet specific rules: The preamble of section 107 specifically recites six uses likely to result in a finding of fair use: “criticism, comment, news reporting, teaching . . ., scholarship, or research.” But determining a “fair use” is “not to be simplified with bright-line rules, for the statute, like the doctrine it recognizes, calls for case-by-case analysis.” Campbell, 510 U.S. at 576-77. As such, the statutory list is not exhaustive. See 17 U.S.C. § 107 identifying purposes “such as” the listed set . The statute’s legislative history confirms the same. See Harper & Row Publishers v. Nation Enters., 471 U.S. 539, 562 1985 ; Pac. & S. Co., Inc. v. Duncan, 744 F.2d 1490, 1496 11th Cir. 1984 ; Cambridge Univ. Press v. Becker, 863 F. Supp. 2d 1190, 1225 N.D. Ga. 2012 . The House Report for section 107 indicates Congress’s intent for a flexible inquiry that can adapt to new technologies and scenarios: T here is no disposition to freeze the doctrine in the statute, especially during a period of rapid technological change. Beyond a very broad statutory explanation of what fair use is and some of the criteria applicable to it, the courts must be free to adapt the doctrine to particular situations on a case-by-case basis. H.R. Rep. No. 94–1476, 94th Cong., 2d Sess. 66 1976 ; see also id. at 65 “ S ince the doctrine is an equitable rule of reason, no generally applicable definition is possible, and each case raising the question must be decided on its own facts.” . The DOJ agrees with Alsup’s analysis that training is quite clearly fair use as transformative. The first statutory fair-use factor, the “purpose and character of the use,” requires consideration of “whether the new work merely ‘supersede s the objects’ of the original creation ‘supplanting’ the original , or instead adds something new, with a further purpose or different character, altering the first with new expression, meaning, or message.” Campbell, 510 U.S. at 578-79 quoting Folsom v. Marsh, 9 F. Cas. 342, 348 C.C.D. Mass. 1841 No. 4,4901 Story, J. and Harper & Row, 471 U.S. at 562 . The Supreme Court has described the latter type of use as “transformative.” Id. “ T ransformative uses tend to favor a fair use finding because a transformative use is one that communicates something new and different from the original or expands its utility, thus serving copyright’s overall objective of contributing to public knowledge.” Authors Guild, 804 F.3d at 214. The copying of protected text articles as part of training an LLM is a use of a different kind or character that is “transformative—spectacularly so.” Bartz v. Anthropic PBC, 787 F. Supp. 3d 1007, 1021 N.D. Cal. 2025 . The New York Times alleges that OpenAI’s training results in a model that “predict s words that are likely to follow a given string of text based on the potentially billions of examples used to train” OpenAI’s LLMs, such that the LLMs can subsequently generate original responses to a wide range of user inputs. Microsoft Corp., No. 1:23-cv-11195-SHS-OTW, ECF 1677 ¶ 75 Aug. 21, 2026 . An OpenAI LLM thus uses the copyrighted work not to duplicate the work’s expressive content, but as part of a process to learn and act on statistical patterns in written text, including vocabulary, syntax, and knowledge. The purpose of the copying to build an intelligent, interactive model differs in kind from the purpose of the copied work to use language to directly entertain or educate a reading audience . This use of text-based works to create an LLM engine for “innovative tools” that can “edit an email . . . , translate an excerpt from or into a foreign language, write a skit based on a hypothetical scenario, or do any number of other tasks” is undoubtedly “highly transformative.” Kadrey v. Meta Platforms, Inc., 788 F. Supp. 3d 1026, 1044 N.D. Cal. 2025 . Notably, the whole concept of “transformative use” being so central to fair use is generally traced back to https://en.wikipedia.org/wiki/Transformative use Leval article Judge Pierre Leval’s wonderful 1990 Law Review article “Toward a Fair Use Standard.” At the time he wrote that, he was a federal judge in the Southern District of NY, where this case is being heard side note: it’s ridiculous that Leval’s “Toward a Fair Use Standard” article seems to mainly only be available behind JSTOR’s paywall https://www.jstor.org/stable/1341457 …. if ever there were an article that should be available freely… . Beyond the transformativeness, the DOJ leans heavily again, correctly on the other big factor that shows up in fair use cases: the impact on the market. As we said when the NY Times first floated this lawsuit https://www.techdirt.com/2023/08/17/ny-times-considering-a-potentially-very-dumb-lawsuit-against-openai-because-it-learned-from-ny-times-content/ , no one is replacing the NY Times with ChatGPT. They serve very different purposes. And the DOJ filing emphasizes this: Using a copyrighted work to train an LLM—without more—generally does not result in this sort of substitution because it does not “reveal ” a significant amount of original “authorial expression.” Authors Guild, 804 F.3d at 224; see also, e.g., Bartz, 787 F. Supp. 3d at 1031 “ T raining LLMs did not result in any exact copies nor even infringing knockoffs of their works being provided to the public.” . In fact, training does not reveal anything to the public at all—it simply creates a copy of a protected work in order to teach an LLM to recognize relationships between data and adapt to new information. The potential for future outputs that might cause market harm is simply not relevant to evaluating an LLM training use under the required use-byuse analysis. To be clear, even LLM outputs that compete with text articles—without reproducing or substantially resembling protected aspects of text articles—would not be substantially similar to, or substitutes for, copyrighted works in the relevant sense. When outputs “copy no protected elements of the original work, much less significant portions,” they cannot cause the relevant form of market harm just because they happen to be “in the same genre or category of works” as the original, given that “a genre is an uncopyrightable idea or method of expression.” Edward Lee, Copyright Dilution Under Constitutional Scrutiny, 25 Chi.-Kent J. Intell. Prop. 1, 6 2026 citing Peters v. West, 692 F.3d 629, 636 7th Cir. 2012 “ N o poet can claim copyright protection in the form of a sonnet or a limerick.” ; accord Abdin v. CBS Broad. Inc., 971 F.3d 57, 70 2d Cir. 2020 no infringement where “an independent comparison of the works reveals that there is no substantial similarity between the protectible features of the original ” and the secondary use . Whether a particular output or category of outputs is substantially similar to the copyrighted work, and whether any substantially similar reproduction might be a significantly competing substitute, are distinct questions involving distinct “challenged use s ” and potentially additional legal questions . Harper & Row, 471 U.S. at 568. But outputs lacking in substantial similarity cannot cause the sort of market harm that is cognizable in the fair-use analysis. I’m also happy to see the DOJ make a point that usually gets ignored in these cases: creative people learn by copying. It is natural. Creative people imitate others until they find their own voice. A ruling that training isn’t fair use would turn that whole creative learning trajectory into infringement: This type of logic would have problematic implications for copyright law generally. T o illustrate, when she was a teenager, Joan Didion “would type out” Ernest Hemingway’s “stories to learn how the sentences worked,” and as a result she considered him the greatest influence on her writing. See Linda Kuehl, Joan Didion, the Art of Fiction No. 71, The Paris Review Issue 74, Fall-Winter 1978 .18 By the Kadrey court’s logic, Didion should have incurred liability to Hemingway every time she published a piece, because the process by which she trained herself and the process by which she produced works was all one use, and her works competed with those of other authors in the market for literature. But “to make anyone pay specifically for the use of a book . . . each time they later draw upon it when writing new things in new ways would be unthinkable.” Bartz, 787 F. Supp. 3d at 1021. I’m also glad to see the DOJ call out the simple fact that, for all of the NY Times’ whining about how awful it is that OpenAI trained on the NY Times and basically every other published work out there , NY Times reporters regularly rely on LLM tools themselves. And… that it’s helping smaller media providers level up to compete with a media company as large and full of resources as the NY Times. LLMs can inspire or help someone to write a story, compose a song, write a movie script, or produce any other kind of art. Indeed, authors at the New York Times itself are using LLMs to help them “conceptualize and edit” articles.19 Independent and start-up publications, as well as ordinary people, can too. For example, an independent writer used an LLM and his background as a physics teacher to offer a contrarian perspective about data center water usage and critique the New York Times. This is a good filing, and it sucks that this DOJ is so untrustworthy that the court may discount it. And while it’s easy to claim that this was just done because of how the tech oligarchs have lined up behind Donald Trump, there are some suggestions that this is not the case here. This filing looks like the work of a few lawyers at the DOJ who actually understand copyright law — which, these days, is its own kind of shocking. And the best sign of this is that Mike Davis, the MAGA whisperer who appears quite gleeful about how https://www.techdirt.com/2026/04/02/wsj-lobbyists-easily-destroyed-any-semi-serious-antitrust-enforcers-left-in-maga/ if you pay him, he’ll get your antitrust case to turn out the way you want, is absolutely freaking out about this filing, and went public with his demand that the DOJ withdraw the filing https://www.foxnews.com/opinion/mike-davis-trump-doj-withdraw-statement-interest-openai-lawsuit in a Fox News op-ed calling it “the art of the steal” — a phrase that reveals he has no idea what fair use is, since a use that isn’t infringement isn’t theft. Also necessary: a reminder that Mike Davis became anti-tech only after big tech companies refused to hire him https://www.politico.com/news/2021/07/22/major-gop-tech-critics-sought-funding-from-google-500602 . The court may decide to ignore it, but the DOJ’s filing is absolutely correct on the issue of fair use. It’s just too bad it’s coming from a DOJ that has spent every last bit of credibility it had on cases that deserved none of it. Filed Under: ai https://www.techdirt.com/tag/ai/ , ai training https://www.techdirt.com/tag/ai-training/ , copyright https://www.techdirt.com/tag/copyright/ , doj https://www.techdirt.com/tag/doj/ , fair use https://www.techdirt.com/tag/fair-use/ , transformative https://www.techdirt.com/tag/transformative/ Companies: ny times https://www.techdirt.com/company/ny-times/ , openai https://www.techdirt.com/company/openai/