{"slug": "nyt-filings-reveal-microsoft-openai-knew-of-theft-risks-in-ai-training", "title": "NYT Filings Reveal Microsoft, OpenAI Knew of ‘Theft’ Risks in AI Training", "summary": "Newly unsealed court filings in The New York Times' copyright lawsuit against OpenAI and Microsoft reveal that Microsoft's director of applied science, Brent Hecht, called large-scale copying of online content for AI training an \"astonishing theft of unprecedented proportions\" in a January 2023 memo, and that Microsoft data shows click-through rates to Times websites from Bing's AI product were 87% to 93% lower than from conventional Bing Search. The filings also detail that between 2019 and 2022 Microsoft supplied OpenAI with Bing Index data under the codename Project Taxi and later launched Project Mango, whose web crawler produced a dataset containing at least 160,903 unique works belonging to the news publishers in the litigation. The Times sued both companies in December 2023, and Microsoft and OpenAI maintain that training on copyrighted material qualifies as fair use under US law.", "body_md": "**September 19, 2026, (Inside AI) —** Newly unsealed court documents in **The New York Times**' copyright lawsuit against **OpenAI** and **Microsoft** reveal that executives at both companies privately acknowledged the legal and ethical risks of training AI models on millions of news articles. The filings, made public this week, include internal memos, deposition transcripts, and company data that could reshape the fair use debate at the heart of the case.\n\nThe documents show that Microsoft's own director of applied science, **Brent Hecht**, described large-scale copying of online content for AI training as an \"astonishing theft of unprecedented proportions\" and perhaps the \"largest theft of labor in human history\" in a January 2023 memo. OpenAI's head of ChatGPT, **Nick Turley**, wrote that the company's products were \"largely substitutive, period,\" predicting they would become more so as they improved. Microsoft CEO **Satya Nadella** acknowledged in a deposition that chatbot answers could replace visits to the underlying website.\n\nThe Times sued both companies in December 2023, alleging unauthorized use of its copyrighted articles to train AI models and that products like ChatGPT could reproduce or closely mimic its journalism. Microsoft and OpenAI have argued that training on copyrighted material qualifies as fair use under US law. The case hinges on whether that training is transformative and whether the resulting products harm the market for the original works.\n\nThe unsealed filings provide the publishers with new ammunition. Microsoft's own data shows that click-through rates to Times websites from Bing's AI product were between **87%** and **93%** lower than from conventional Bing Search. Other publisher groups cited in the litigation saw similar declines. These figures do not by themselves prove infringement, but they support the argument that AI products substitute for and damage the market for the journalism used to build them.\n\n**Read:** **ACC Sues The L Suite for Training Chatbot on Copyrighted Legal Materials**\n\n\"Microsoft's position is set out in its court filings, which explain why these transformative uses are consistent with copyright law and why Copilot is not a substitute for publishers' journalism,\" a Microsoft spokesman said. A separate company filing described Hecht as a research academic who also worked at **Northwestern University**, and said he was employed by Microsoft \"to present divergent and asymmetric perspectives,\" not to be a decision-maker.\n\nThe filings also detail how the companies acquired training data. Between 2019 and 2022, Microsoft supplied OpenAI with data from its Bing Index, a compilation of billions of webpages, under a codename **Project Taxi**. The transfer included Times content, which Microsoft sold to OpenAI in an undisclosed transaction. Microsoft later launched **Project Mango** to help OpenAI \"collect as many of the documents as possible for...training the large language models.\" The Mango web crawler copied web content on OpenAI's behalf, producing a dataset containing at least **160,903** unique works belonging to the news publishers involved in the litigation.\n\nThe documents also contain allegations that OpenAI employees found ways to access paywalled material. In one exchange, an OpenAI researcher told company president **Greg Brockman** about a \"hack to get around\" the Times paywall; Brockman replied, \"ah nice\". In a deposition, Nadella said that \"anything that is paywalled should be licensed by anyone who wants to use it.\" He also said that if he \"had been made aware that OpenAI had scraped and trained on information that was behind a paywall,\" he would have invoked Microsoft's right to require OpenAI to retrain its models.\n\nThe legal question remains unresolved. An employee's description of copying as \"theft\" does not determine whether it constitutes copyright infringement. Courts weigh four statutory factors: the purpose and character of the use, the nature of the copyrighted work, the amount used, and the effect on the potential market. The judge will decide whether the use was fair based on the circumstances.\n\n**Read:** **DOJ Will Probe AI-Related Crimes, Attorney General Blanche Says**\n\nFor users, the issue is straightforward: AI tools can increasingly answer questions without requiring visits to the websites where that information originated. The filings show that OpenAI and Microsoft were aware of what this could mean for the publishers whose work these systems rely on. The case continues to test how US copyright law applies to generative AI, with implications for the entire news industry and the future of AI training data.", "url": "https://wpnews.pro/news/nyt-filings-reveal-microsoft-openai-knew-of-theft-risks-in-ai-training", "canonical_source": "https://insideai.news/news/ai-policy-and-regulation/nyt-openai-microsoft-lawsuit/12346/", "published_at": "2026-09-19 13:06:55+00:00", "updated_at": "2026-09-19 13:25:49.211209+00:00", "lang": "en", "topics": ["ai-policy", "ai-ethics", "artificial-intelligence", "large-language-models", "ai-crawlers"], "entities": ["The New York Times", "OpenAI", "Microsoft", "Brent Hecht", "Nick Turley", "Satya Nadella", "Greg Brockman", "Project Mango"], "alternates": {"html": "https://wpnews.pro/news/nyt-filings-reveal-microsoft-openai-knew-of-theft-risks-in-ai-training", "markdown": "https://wpnews.pro/news/nyt-filings-reveal-microsoft-openai-knew-of-theft-risks-in-ai-training.md", "text": "https://wpnews.pro/news/nyt-filings-reveal-microsoft-openai-knew-of-theft-risks-in-ai-training.txt", "jsonld": "https://wpnews.pro/news/nyt-filings-reveal-microsoft-openai-knew-of-theft-risks-in-ai-training.jsonld"}}