OpenAI knew it was training its AI on pirated books, its own messages show Unsealed filings published by the Authors Guild on September 21 in Authors Guild v. OpenAI and Microsoft show OpenAI staff knew early models were trained on books from the pirate library LibGen and expected the practice to hurt authors, with one researcher writing he was "just worried about optics." The filings state OpenAI deleted its LibGen files in June 2022 under an effort called Project Clear, and that Microsoft knew about the pirated data from April 2019 when Sam Altman and Dario Amodei showed an early GPT-3 to Bill Gates and Microsoft CTO Kevin Scott. Authors including John Grisham, Jodi Picoult, Jonathan Franzen and George R.R. Martin filed for summary judgment on September 4, while OpenAI has moved for fair use; the court has not ruled on either motion. Fiction on the shelves at Greenlight Bookstore in Brooklyn, New York illustrative . Image: Sashimi-b https://commons.wikimedia.org/wiki/File:Greenlight Bookstore interior 2026.jpg / Wikimedia Commons, CC BY 4.0 https://creativecommons.org/licenses/by/4.0/ , cropped OpenAI’s own staff knew the company was training its AI on books copied from a piracy website, and expected it to hurt authors, according to internal messages made public in the authors’ copyright lawsuit against OpenAI and Microsoft. The court filings were published by the Authors Guild https://authorsguild.org/news/ag-v-openai-top-execs-knew-mass-book-piracy-was-illegal/ on September 21. They’re getting fresh attention today for an ironic reason. One OpenAI researcher had worried that news of the pirated books would end up on Hacker News, the popular tech forum. On Sunday morning, the filings went to No. 1 on Hacker News https://news.ycombinator.com/item?id=49863864 . What the filings say The documents come from briefs in Authors Guild v. OpenAI and Microsoft , part of a larger set of copyright cases combined in federal court in Manhattan. According to the Authors Guild, OpenAI trained early models on books from LibGen, a huge pirate library, and Microsoft knew about it from April 2019, when Sam Altman and Dario Amodei showed an early version of GPT-3 to Bill Gates and Microsoft CTO Kevin Scott. The worry inside OpenAI, at least for one researcher, seems to have been less about the law than about how it would look. He wrote that he was “just worried about optics”: ‘openai uses copyrighted data from sketchy russian website’ showing up on Hacker News would be unfortunate. Sam McCandlish, then an OpenAI researcher, in the unsealed filings Amodei, then OpenAI’s research director, is quoted saying “as a training set LibGen is a bit sketchier.” Publishers Weekly https://www.publishersweekly.com/pw/by-topic/digital/copyright/article/101300-unsealed-files-show-open-ai-microsoft-knew-copying-was-illegal-and-could-hurt-authors.html also cites an internal note from August 2019: “We trained GPT-3 on pirated stuff No sharing that ” “Project Clear” In June 2022 OpenAI deleted its LibGen files, in an effort the filings call Project Clear. Bob McGrew, then OpenAI’s VP of research, wrote that “given how much OpenAI is in the news, now is the right time to excise Libgen from our systems and storage.” Earlier in the case, OpenAI was ordered to hand over internal Slack channels named project-clear and excise-libgen and to make its in-house lawyers available to be questioned about why the data was deleted. “We’ll likely ignore their concerns” The filings also show staff openly expected AI to hurt writers. Jack Clark, then OpenAI’s policy director, wrote in May 2020: There will be a point where a bunch of artists express worry about what we’re doing here and we’ll likely ignore their concerns and release anyway. Jack Clark, then OpenAI’s policy director In the same message he said “the better we do on GPT-X, the more worried genre fiction authors will become about us substituting for them on Amazon.” Another employee, Tarun Gogineni, wrote that “even if GRRM dies early, GPT-5 will autocomplete his series,” a reference to George R.R. Martin, who is one of the plaintiffs, and described authors’ complaints as “acceptable economic disruption.” Amodei and Clark later left OpenAI and co-founded Anthropic in 2021, which has faced the same problem: it agreed to pay $1.5 billion in 2025 to settle a lawsuit by authors over pirated books it had downloaded. What happens next The authors, who include John Grisham, Jodi Picoult, Jonathan Franzen and Martin, filed for summary judgment on September 4, asking the court to rule in their favour without a full trial. OpenAI has filed its own motion arguing that training AI on books is fair use and that its models rarely reproduce authors’ work. The court hasn’t ruled on either. OpenAI’s appetite for books hasn’t gone away. On Saturday, the Guardian revealed that Oxford had let OpenAI train its AI on texts digitised from the Bodleian Library https://madrobot.blog/2026/09/26/oxford-bodleian-library-openai-training-data-chatgpt/ . Authors Guild CEO Mary Rasenberger said the filings reveal “shocking disdain for writers and their work.” Why it matters AI companies have long argued that training on books is fair use. These documents don’t settle that legal question, but they show OpenAI staff knew the books were pirated, expected the technology to replace authors and deleted the evidence as scrutiny grew, which could weigh heavily with a court deciding whether the copying was fair. Sources: Authors Guild https://authorsguild.org/news/ag-v-openai-top-execs-knew-mass-book-piracy-was-illegal/ , Publishers Weekly https://www.publishersweekly.com/pw/by-topic/digital/copyright/article/101300-unsealed-files-show-open-ai-microsoft-knew-copying-was-illegal-and-could-hurt-authors.html , The Hollywood Reporter https://www.hollywoodreporter.com/business/business-news/openai-loses-key-discovery-battle-why-deleted-library-of-pirated-books-1236436363/ , Hacker News https://news.ycombinator.com/item?id=49863864 .