# Training AI on copyrighted books might actually be legal under

> Source: <https://promptcube3.com/en/news/7457/>
> Published: 2026-08-24 01:22:04+00:00

# Training AI on copyrighted books might actually be legal under

The entire debate hinges on whether AI training constitutes "Fair Use" under US copyright law. If you look at the core of the argument, AI companies aren't typically republishing the books; they are analyzing the patterns, structures, and relationships between words to build a mathematical model. They argue this is "transformative"—meaning they are creating something entirely new from the raw data, much like how a human student reads a thousand books to learn how to write a novel.

## The core arguments for Fair Use

Proponents of AI development rely on several key pillars to justify their data scraping practices:

**Transformative Purpose:** The goal isn't to provide a free version of the book to the user, but to teach a machine the mechanics of language.**Non-Expressive Use:** This is a technical distinction. The AI isn't "copying" the creative expression (the story or the soul of the book); it is "ingesting" the data points to understand statistical probability.**Market Impact:** Companies argue that an LLM isn't a direct substitute for a specific novel, so it doesn't technically harm the market for that individual book in the same way a pirated PDF would.

## Why authors are fighting back

The counter-argument from the creative community is gaining massive momentum, and for good reason. If a model can be prompted to "write a story in the exact style of [Author Name]," the model has effectively commodified that author's lifelong work without compensation.

The legal battleground is shifting toward "output" rather than just "input." While the training phase might be defended as transformative, the ability of an LLM agent to mimic a specific creator's unique voice or style starts to look a lot more like copyright infringement. We are seeing a move toward a more structured AI workflow where developers might eventually be forced to license datasets through official channels, similar to how Spotify pays music labels.

If the courts decide that training is not Fair Use, the entire industry faces a massive hurdle. We would likely see a shift toward "permission-based" datasets, where high-quality, licensed text becomes the gold standard for fine-tuning models. This would move us away from the current "scrape everything" mentality and toward a more sustainable, albeit more expensive, deployment model for foundational models. For now, we are stuck in a period of intense litigation that will ultimately define the boundaries of prompt engineering and data ownership for decades to come.

[British Gen Z trusts AI less than Boomers — here's the data 2d ago](/en/news/7195/)

[Berkeley Law just banned AI by default across every classroom 2d ago](/en/news/7131/)

[EU copyright office confirms AI output falls outside protection 3d ago](/en/news/7124/)

[Twitch is now using your stream data to train generative AI by 10d ago](/en/news/6161/)

[Generative AI is basically the Guitar Hero of the creative world 16d ago](/en/news/5421/)

[AI-Designed Bacteriophages: Engineering a Novel E. coli Killer 17d ago](/en/news/5330/)

[Next Linkdaze is actually building a household OS instead of a basic →](/en/news/7455/)
