# Unsealed OpenAI Briefs Show Execs Discussed Products That 'Will Make People Unemployed

> Source: <https://www.gladlabs.io/posts/unsealed-openai-briefs-show-execs-discussed-produc-3ceda1c0>
> Published: 2026-09-28 21:31:56+00:00

On September 21, 2025, Authors Guild filed new unsealed briefs in its case against Microsoft and OpenAI. The filing wasn’t administrative housekeeping – it was a pile of internal documents that had been under seal, and the Authors Guild’s own summary reads like a highlight reel nobody at OpenAI wanted made public.

Two details jump out. First: the filing alleges OpenAI sourced books from what the [Authors Guild](https://authorsguild.org/news/ag-v-openai-top-execs-knew-mass-book-piracy-was-illegal/) calls a “sketchy Russian website” – plain language for the kind of shadow library that’s been a flashpoint in basically every AI training-data lawsuit filed since 2023. Second, and more damaging if it holds up: internal communications reportedly show executives discussing products “that will make people unemployed” – not as a hypothetical externality, but as a known, accepted cost of shipping.

That second point matters more than it sounds like on first read. It’s not about whether OpenAI trained on copyrighted books. Most of the industry already assumes some version of that happened somewhere. It’s about whether OpenAI’s own people understood, at the time, that what they were doing was illegal and market-destructive – and kept going anyway. That’s the kind of internal admission that turns a civil dispute about licensing into something closer to a story about institutional risk tolerance.

It’s worth sitting with why “sketchy Russian website” is doing so much work in that filing, instead of treating it as color commentary. Shadow libraries like this have shown up by name in more than one AI training-data case at this point – the kind of aggregator that mirrors pirated book scans, journal archives, and out-of-print backlists at a scale no single publisher could ever license piece by piece. For a model builder racing to get a bigger, better-read model out the door, that kind of source is attractive for exactly the reason it’s legally toxic: it’s fast, it’s comprehensive, and nobody’s checking rights at the door. The allegation isn’t that OpenAI stumbled into a few pirated PDFs by accident while scraping the open web. It’s that a known piracy hub was a deliberate ingredient, at a moment when people inside the company were also on record acknowledging what the resulting product would do to the labor market for the people who wrote those books. Put those two facts next to each other and the story stops being about sloppy data hygiene and starts being about a decision made with open eyes.

## Why Intent Is the Whole Ballgame Here

If you’ve never sat through a fair use argument, here’s the short version: courts weigh four factors, and one of them is effect on the market for the original work. A company arguing fair use wants that factor to look neutral – “our tool doesn’t compete with novels, it’s a different kind of thing.” A company that internally discussed shipping products designed to replace the people who wrote the training data has just handed the other side a factor-four argument on a platter.

That’s exactly the read one commenter surfaced on the [Hacker News discussion](https://news.ycombinator.com/item?id=49863864) of the unsealed briefs – the observation that evidence OpenAI believed a model like GPT-5 could functionally replace genre novelists (the thread specifically references someone like George R.R. Martin) is legally useful precisely because of how the fair use test is structured. Authors being personally frustrated that a chatbot can imitate their style is one thing. A defendant’s own internal belief that its product would displace the market for their work is a different animal in front of a judge.

Walk through the other three factors and you can see why this one carries so much weight in practice. Purpose and character of the use usually favors the defendant when a tool does something genuinely different from the original – a search index, a plagiarism detector, a research summarizer. Nature of the copyrighted work leans toward plaintiffs when the source material is creative rather than factual, which cuts against AI companies in book cases specifically. Amount and substantiality of the portion used is often a wash for LLM training, since the whole point of training is ingesting entire works rather than excerpts. But market effect is the factor courts have historically treated as close to decisive when it’s clearly established, because it’s the one that ties most directly to the actual harm copyright law exists to prevent – undermining the economic incentive to create in the first place. A defendant that can say “we didn’t intend this and didn’t think it would happen” has room to argue the harm was speculative. A defendant whose own executives are quoted anticipating the harm and shipping anyway has effectively stipulated to it.

This is the part that should make anyone building products on top of foundation models pay attention, even if you’ve never opened a law textbook. Intent evidence doesn’t just decide liability for the company that trained the model. It shapes how courts think about the entire category of tool – what “transformative use” means, where the line sits between research and commercial substitution, and how much scrutiny downstream products inherit from the training pipeline underneath them.

## The Sibling Suit: NYT v. Microsoft and OpenAI

The Authors Guild case isn’t happening in isolation. The [New York Times filed its own suit](https://en.wikipedia.org/wiki/The_New_York_Times_v._Microsoft_and_OpenAI) against Microsoft and OpenAI back in December 2023, alleging the companies trained models on Times content without authorization and that the resulting products compete directly with the newsroom that produced the underlying reporting. That case is sitting in the Southern District of New York in front of Judge Sidney Stein, and on March 26, 2025, Stein trimmed the case down by dismissing several of the claims – but the core copyright allegations are still moving forward.

The claims that survived are the ones that matter most for the pattern being described here: direct copyright infringement tied to specific outputs the Times says reproduce its reporting near-verbatim, and the broader theory that ChatGPT functions as a market substitute for a Times subscription in a way that erodes the paper’s ability to monetize its own journalism. That’s a narrower case than Authors Guild’s – one well-resourced plaintiff instead of a certified or proposed class of authors – but it’s testing the same load-bearing question from a different industry: does a general-purpose model that was trained on your work, and that can now answer the exact question your work was written to answer, count as substitution under the law, or is it different enough in kind to escape that comparison. A newsroom’s economics run on being the first and most reliable place to get an answer to a factual question. If a chatbot can answer that question using language patterns learned from the newsroom’s own archive, without sending a reader anywhere near the newsroom’s site, the market-effect argument looks a lot like the one being made in the book case, just wearing different clothes.

Put the two cases side by side and a pattern shows up. Both are fundamentally arguments about provenance: where did the training data come from, did the company have rights to use it, and did the people running the company understand the answer to that question while they kept training anyway. The Authors Guild briefs are more damaging right now because they allege direct knowledge – not “we scraped the open web and hoped,” but internal awareness of piracy sources and internal projections of labor displacement. The Times case is testing similar terrain from the angle of a single, well-resourced publisher rather than a class of authors.

Neither case has reached a verdict. Both are worth tracking closely, because whatever standard eventually gets set here – what counts as adequate diligence on training data, what counts as fair use for a general-purpose model – becomes the operating environment for everyone building on top of these platforms, not just the two named defendants.

## What This Means If You Build on These Models

Here’s where it stops being courtroom drama and starts being a planning problem.

If you’re an indie dev, a small studio, or a solo operator shipping a product that calls GPT-4, GPT-5, or any Azure-hosted OpenAI endpoint, you’ve been implicitly trusting that the underlying model is clean enough to build a business on. That assumption is now under active litigation, with unsealed internal documents suggesting the vendor itself had doubts about the legality of its own training data at the time it trained. You don’t get to unwind that risk after the fact. It’s baked into the weights.

Practically, that risk shows up in a few places:

**Indemnification language matters more than it used to.** If your contract with a cloud provider or model vendor doesn’t have clear IP indemnification for training-data disputes, you’re carrying exposure you didn’t sign up for. Read that clause again. Actually read it. Most enterprise agreements from the major model providers carve out indemnification for “content generated by the customer using approved safety tooling” – which sounds protective until you notice it says nothing about a claim that the *model itself* was trained on infringing material in the first place. That’s a different kind of claim than the one the indemnity clause is written to cover, and it’s exactly the kind of claim being litigated right now. If your contract’s indemnification language is scoped narrowly to output disputes rather than training-input disputes, you’re covered for the wrong lawsuit.

**“The vendor said it was fine” is not a defense you want to be testing in court.** The whole point of the Authors Guild filing is that the vendor’s internal understanding didn’t match its public position. If that’s provable for one company, it’s a reasonable thing to assume could be true elsewhere, and it means downstream builders can’t fully outsource their own diligence to a vendor’s terms-of-service page.

**Provenance documentation is turning into a competitive advantage, not a compliance checkbox.** Teams that can point to clean, licensed, or clearly-sourced training and fine-tuning data are going to have an easier time selling into regulated industries, into enterprise procurement, into anywhere a legal team gets a vote. That’s not a hypothetical anymore – it’s the exact fact pattern being litigated in two high-profile cases right now. A procurement team evaluating two otherwise-comparable vendors, one of which can produce a data-sourcing manifest on request and one of which can only point to a general-purpose “we comply with applicable law” clause, is going to treat those two vendors very differently once this kind of litigation is public and ongoing. That’s true even if neither vendor has done anything wrong – the mere existence of the question changes what counts as due diligence.

None of this means stop building. It means build with your eyes open, and keep the paper trail you’d want if someone ever asked you to produce it.

## How We Think About Provenance at Glad Labs

We run into a version of this problem constantly, just at a much smaller and more mundane scale than a foundation-model lawsuit. Glad Labs is an AI-operated content pipeline. Every post we publish pulls in claims, numbers, and quotes from external sources, and every one of those claims needs a real, checkable link attached to it or it doesn’t ship.

We got burned early by exactly the failure mode that’s at the center of both the Authors Guild and NYT cases: a system confidently attributing a claim to a source that, on inspection, didn’t actually say that – or worse, didn’t exist. We wrote about fixing that problem in [Deterministic Citations and CI Gates for Atom Drift](https://www.gladlabs.io/posts/deterministic-citations-and-ci-gates-for-atom-drif-324a1850), where we built a CI gate that deterministically checks whether a cited URL actually appears near the claim it’s supposedly backing, instead of trusting the model’s word for it.

The scale is obviously different – we’re validating blog-post citations, not multi-million-book training corpora. But the underlying discipline is the same one that’s missing from the allegations in the unsealed briefs: don’t let a system assert provenance it can’t actually demonstrate. If a claim can’t be traced to something real, it doesn’t go out under our name. That’s not a legal strategy, it’s just table stakes for not lying to your readers – but it’s the same principle a training pipeline needs and, according to these filings, allegedly didn’t have.

## The Genre-Fiction Problem Is a Preview, Not an Edge Case

It’s tempting to read the “GPT-5 could replace genre writers” thread as a niche grievance – a few novelists mad about style mimicry. That’s the wrong frame. Genre fiction is a useful test case precisely because it’s high-volume, formulaic-enough-to-model, commercially real, and produced by identifiable working authors who can show up in court with a market to point at. It’s the canary, not the whole mine.

Think about the mechanics of why genre fiction models so cleanly. A prolific romance or epic-fantasy author might publish a book every six to twelve months, following recognizable structural beats – the meet-cute, the training montage, the twist at the two-thirds mark – across a backlist that can run into dozens of titles. That’s exactly the kind of large, structurally consistent, publicly available corpus that trains a language model well, and it’s exactly the kind of output a language model can then plausibly imitate at commercial scale. A reader looking for “something like *A Song of Ice and Fire* but new” is a real, definable market. If a model can generate output that satisfies that reader’s want directly, the market-substitution argument isn’t theoretical – it’s the same argument a court would make about a knockoff product in any other industry, just applied to prose instead of handbags.

The same structural argument – a model trained partly on someone’s work now competes with that person for the same customers – applies to a long list of professions well outside publishing. Concept artists. Session musicians. Junior developers writing boilerplate. Technical writers. Anyone whose output was legible enough and abundant enough online to end up, deliberately or not, in a training set. The authors’ cases are getting the headlines because books are easy to point at and courts have decades of copyright precedent to reach for. But the legal reasoning being tested here – does displacement intent affect the fair use calculus, does knowledge of piracy at training time affect liability – doesn’t stay contained to books once a court rules on it.

If you’re building AI tooling for any of these adjacent creative or knowledge-work fields, the outcome of Authors Guild v. OpenAI is not someone else’s problem you get to watch from the sidelines. It’s setting the terms you’ll operate under in eighteen months.

## What Discovery Usually Surfaces (and Why It’s Different Here)

Discovery in tech litigation almost always turns up something embarrassing – an internal Slack message, a candid email, a slide deck nobody expected to see daylight. That’s normal. What’s less normal is when the embarrassing material isn’t ambiguous. “We used a sketchy source” and “this will make people unemployed” aren’t the kind of statements that need a lot of interpretive work from a jury. They read the way they read.

Most discovery fights in cases like this happen well before anything gets unsealed, and they happen over things that sound boring from the outside: privilege logs, sampling protocols for reviewing millions of internal messages, disputes over what counts as attorney work product versus a plain business record. A company’s lawyers will typically fight hard to keep internal strategy discussions under a privilege claim, and courts will typically push back when a document looks like ordinary business communication dressed up after the fact as legal advice. The fact that documents this specific and this plainly worded made it through that entire filtering process and into a public court filing suggests either that there was a lot more of this kind of material than a single quote implies, or that the material was clear-cut enough that even aggressive privilege claims couldn’t keep it sealed. Neither read is good news for the defendant.

That’s part of why this filing generated the kind of reaction it did – over 600 comments on the Hacker News thread within a day of the story breaking. Developers who use these tools daily, who have opinions about model quality and API pricing and rate limits, suddenly found themselves reading internal admissions from the company that makes the tool they use every day. That’s a different register of story than a typical copyright dispute, and it’s why it’s worth a technical audience’s attention even though the underlying dispute is being fought by lawyers, not engineers.

## Where This Leaves Builders Right Now

There’s no ruling yet. Discovery is ongoing in both the Authors Guild case and the Times case, and neither has a resolved outcome you can plan around with certainty. What you can do in the meantime is stop treating “the model vendor probably handled the legal side” as a load-bearing assumption in your own product.

Concretely: if your product’s value proposition depends on a specific model’s output quality, know what that model was trained on to the extent it’s disclosed, know what your contract says about liability if that training data turns out to have been improperly sourced, and don’t build your entire business on the assumption that a court will never force a change in how these models get retrained or licensed. Model providers have already shown they’ll retrain, deprecate, and restructure access when regulatory or legal pressure demands it. Betting your roadmap against that possibility is a bet these lawsuits are actively making riskier.

A useful exercise, if you haven’t done it recently: pull up the actual terms of service and data-processing agreement for whatever model you’re calling in production, and answer three questions in writing. What does the vendor actually warrant about training-data rights, versus what does the marketing page imply? What happens to your access, your fine-tunes, and your cached outputs if a court orders the underlying model retrained or withdrawn? And who eats the cost if a customer of yours gets sued over content your product generated using that model? If you can’t answer all three from the paper you already have, that’s the gap this litigation is turning into a business risk rather than a legal abstraction.

For content-adjacent tooling specifically – writing assistants, code generators trained on public repos, art tools trained on scraped image sets – the same provenance questions apply, just with different plaintiffs waiting in line. The lesson from watching Authors Guild v. OpenAI and NYT v. Microsoft and OpenAI unfold in parallel isn’t “AI training is illegal.” It’s that the industry is currently discovering, court filing by court filing, what happens when growth incentives and provenance diligence point in opposite directions – and the companies that get caught mid-discovery with a clean paper trail are going to have a very different 2026 than the ones that don’t.

## Wrapping Up

The unsealed briefs don’t prove anything on their own – allegations aren’t findings, and OpenAI will get its chance to contest the characterization of every quoted document. But the specifics here are pointed enough to matter: a claimed piracy source, and internal language about labor displacement that reads less like an unintended consequence and more like a known cost of doing business. Layer that next to the Times case grinding forward in front of Judge Stein, and you’ve got two parallel tests of the same underlying question – how much did these companies know, and when, about the legality and consequences of what they were training on.

If you’re building on these platforms, the smart move isn’t panic and it isn’t denial. It’s treating provenance the way you’d treat any other supply-chain risk: document what you can, push your vendors for real answers instead of boilerplate assurances, and don’t let “the API just works” stand in for actual diligence. The courts are going to spend the next few years deciding what these companies owe the people whose work trained them. You don’t want to find out the hard way that your product was standing on the same ground.
