Newly unsealed court files show that what AI companies say behind the scenes doesn’t always line up with what they tell the public.
In a legal brief unsealed this week, an executive at Microsoft described OpenAI’s choice to train its models on news articles and other human-published content as the “largest theft of labor in human history.” The brief came to light through a copyright lawsuit the New York Times filed against Microsoft and OpenAI in 2023 over their practice of using the paper’s news stories to train artificial intelligence.
Microsoft’s Director of Applied Science Brent Hecht made the statement, according to the brief, and also called the process of training AI models on copyrighted works “an astonishing theft of unprecedented proportions.” The brief also revealed that OpenAI’s head of ChatGPT admitted that publishers like the New York Times “face an existential threat” from AI products, which he described as “largely substitutive, period” – that is, designed to replace their source material. Hecht also described large language models as “a product that destroys its supply chain.”
In its lawsuit, the New York Times names Copilot and ChatGPT specifically, noting that the products were created by “copying and using millions of The Times’s copyrighted news articles, in-depth investigations, opinion pieces, reviews, how-to guides, and more.”
The newspaper argues that by mass-ingesting its news stories without permission in order to train large language models designed to replace newsrooms, the pair of AI companies violated copyright law. In the suit, the publisher accuses Microsoft and Open AI seeking a “free ride,” leveraging the paper’s own monetary investment into quality journalism to “build substitutive products without permission or payment.”
“While Defendants engaged in widescale copying from many sources, they gave Times content particular emphasis when building their LLMs—revealing a preference that recognizes the value of those works,” the paper’s legal team stated in its suit.
In the years since the New York Times first sued the two companies, parallel lawsuits from the New York Daily News and the Center for Investigative Reporting have been combined into one consolidated effort that hopes to hold AI companies to account. Other plaintiffs on the mega-suit now include the Intercept, Ziff Davis, and local newspapers like Chicago Tribune and the Orlando Sentinel.
Executives at both Microsoft and OpenAI have maintained that the business of powering their lucrative AI systems with news stories falls under fair use, a legal framework that protects “transformative” copies of copyrighted works. The definition of that term leaves plenty of room for interpretation, a tension that lies at the heart of the debate over generative AI’s existential dependence on copyrighted art, writing, and other human-crafted original material.
In the new brief, the Times points to the AI executives’ private statements as evidence that their products were designed from the jump as a substitute for the material they were trained on, an argument against the fair use defense. The internal statements come from documents surfaced during discovery and from statements made in depositions with company executives.
“Resolution of these issues for Plaintiffs as a matter of law will streamline this case for trial and, by rejecting Defendants’ efforts to recast free riding as fair use, recognize that the future not just of journalism but of responsible AI too depends on preserving incentives for humans to produce the creative works on which a healthy society depends,” the brief states.