{"slug": "openai-knew-it-was-training-its-ai-on-pirated-books-its-own-messages-show", "title": "OpenAI knew it was training its AI on pirated books, its own messages show", "summary": "Unsealed filings published by the Authors Guild on September 21 in Authors Guild v. OpenAI and Microsoft show OpenAI staff knew early models were trained on books from the pirate library LibGen and expected the practice to hurt authors, with one researcher writing he was \"just worried about optics.\" The filings state OpenAI deleted its LibGen files in June 2022 under an effort called Project Clear, and that Microsoft knew about the pirated data from April 2019 when Sam Altman and Dario Amodei showed an early GPT-3 to Bill Gates and Microsoft CTO Kevin Scott. Authors including John Grisham, Jodi Picoult, Jonathan Franzen and George R.R. Martin filed for summary judgment on September 4, while OpenAI has moved for fair use; the court has not ruled on either motion.", "body_md": "Fiction on the shelves at Greenlight Bookstore in Brooklyn, New York (illustrative). Image: [Sashimi-b](https://commons.wikimedia.org/wiki/File:Greenlight_Bookstore_interior_2026.jpg) / Wikimedia Commons, [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/), cropped\n\nOpenAI’s own staff knew the company was training its AI on books copied from a piracy website, and expected it to hurt authors, according to internal messages made public in the authors’ copyright lawsuit against OpenAI and Microsoft. The court filings were [published by the Authors Guild](https://authorsguild.org/news/ag-v-openai-top-execs-knew-mass-book-piracy-was-illegal/) on September 21.\n\nThey’re getting fresh attention today for an ironic reason. One OpenAI researcher had worried that news of the pirated books would end up on Hacker News, the popular tech forum. On Sunday morning, the filings went to [No. 1 on Hacker News](https://news.ycombinator.com/item?id=49863864).\n\n## What the filings say\n\nThe documents come from briefs in *Authors Guild v. OpenAI and Microsoft*, part of a larger set of copyright cases combined in federal court in Manhattan. According to the Authors Guild, OpenAI trained early models on books from LibGen, a huge pirate library, and Microsoft knew about it from April 2019, when Sam Altman and Dario Amodei showed an early version of GPT-3 to Bill Gates and Microsoft CTO Kevin Scott.\n\nThe worry inside OpenAI, at least for one researcher, seems to have been less about the law than about how it would look. He wrote that he was “just worried about optics”:\n\n‘openai uses copyrighted data from sketchy russian website’ showing up on Hacker News would be unfortunate.\n\nSam McCandlish, then an OpenAI researcher, in the unsealed filings\n\nAmodei, then OpenAI’s research director, is quoted saying “as a training set LibGen is a bit sketchier.” [Publishers Weekly](https://www.publishersweekly.com/pw/by-topic/digital/copyright/article/101300-unsealed-files-show-open-ai-microsoft-knew-copying-was-illegal-and-could-hurt-authors.html) also cites an internal note from August 2019: “We trained GPT-3 on pirated stuff! No sharing that!”\n\n## “Project Clear”\n\nIn June 2022 OpenAI deleted its LibGen files, in an effort the filings call Project Clear. Bob McGrew, then OpenAI’s VP of research, wrote that “given how much OpenAI is in the news, now is the right time to excise Libgen from our systems and storage.” Earlier in the case, OpenAI was ordered to hand over internal Slack channels named project-clear and excise-libgen and to make its in-house lawyers available to be questioned about why the data was deleted.\n\n## “We’ll likely ignore their concerns”\n\nThe filings also show staff openly expected AI to hurt writers. Jack Clark, then OpenAI’s policy director, wrote in May 2020:\n\nThere will be a point where a bunch of artists express worry about what we’re doing here and we’ll likely ignore their concerns and release anyway.\n\nJack Clark, then OpenAI’s policy director\n\nIn the same message he said “the better we do on GPT-X, the more worried genre fiction authors will become about us substituting for them on Amazon.” Another employee, Tarun Gogineni, wrote that “even if GRRM dies early, GPT-5 will autocomplete his series,” a reference to George R.R. Martin, who is one of the plaintiffs, and described authors’ complaints as “acceptable economic disruption.”\n\nAmodei and Clark later left OpenAI and co-founded Anthropic in 2021, which has faced the same problem: it agreed to pay $1.5 billion in 2025 to settle a lawsuit by authors over pirated books it had downloaded.\n\n## What happens next\n\nThe authors, who include John Grisham, Jodi Picoult, Jonathan Franzen and Martin, filed for summary judgment on September 4, asking the court to rule in their favour without a full trial. OpenAI has filed its own motion arguing that training AI on books is fair use and that its models rarely reproduce authors’ work. The court hasn’t ruled on either.\n\nOpenAI’s appetite for books hasn’t gone away. On Saturday, the Guardian revealed that [Oxford had let OpenAI train its AI on texts digitised from the Bodleian Library](https://madrobot.blog/2026/09/26/oxford-bodleian-library-openai-training-data-chatgpt/).\n\nAuthors Guild CEO Mary Rasenberger said the filings reveal “shocking disdain for writers and their work.”\n\n## Why it matters\n\nAI companies have long argued that training on books is fair use. These documents don’t settle that legal question, but they show OpenAI staff knew the books were pirated, expected the technology to replace authors and deleted the evidence as scrutiny grew, which could weigh heavily with a court deciding whether the copying was fair.\n\n*Sources: [Authors Guild](https://authorsguild.org/news/ag-v-openai-top-execs-knew-mass-book-piracy-was-illegal/), [Publishers Weekly](https://www.publishersweekly.com/pw/by-topic/digital/copyright/article/101300-unsealed-files-show-open-ai-microsoft-knew-copying-was-illegal-and-could-hurt-authors.html), [The Hollywood Reporter](https://www.hollywoodreporter.com/business/business-news/openai-loses-key-discovery-battle-why-deleted-library-of-pirated-books-1236436363/), [Hacker News](https://news.ycombinator.com/item?id=49863864).*", "url": "https://wpnews.pro/news/openai-knew-it-was-training-its-ai-on-pirated-books-its-own-messages-show", "canonical_source": "https://madrobot.blog/2026/09/27/openai-libgen-pirated-books-hacker-news-optics-authors-guild-filings/", "published_at": "2026-09-27 09:06:00+00:00", "updated_at": "2026-09-27 09:58:42.760293+00:00", "lang": "en", "topics": ["ai-policy", "large-language-models", "artificial-intelligence", "ai-ethics"], "entities": ["OpenAI", "Microsoft", "Authors Guild", "LibGen", "Sam Altman", "Dario Amodei", "Jack Clark", "Anthropic"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/openai-knew-it-was-training-its-ai-on-pirated-books-its-own-messages-show", "markdown": "https://wpnews.pro/news/openai-knew-it-was-training-its-ai-on-pirated-books-its-own-messages-show.md", "text": "https://wpnews.pro/news/openai-knew-it-was-training-its-ai-on-pirated-books-its-own-messages-show.txt", "jsonld": "https://wpnews.pro/news/openai-knew-it-was-training-its-ai-on-pirated-books-its-own-messages-show.jsonld"}}