Amazon, which started off selling books, is destroying rare texts to train AI
Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.
Amazon is now physically destroying rare, out-of-print books to feed their LLM training pipelines. This unlocks a vast, uncontaminated corpus of pre-2022 human text that can prevent model collapse and improve output quality, but it also means any team relying on public-domain or licensed data is now competing with a proprietary, non-reproducible dataset that only Amazon can build.
Amazon is systematically purchasing and destroying rare, out-of-print books to digitize them for LLM training at a dedicated physical scanning facility. This aggressive offline ingestion highlights the critical shortage of clean, pre-2022 human text, signaling to production teams that the high-quality, non-synthetic data required to prevent model collapse is increasingly gatekept behind proprietary physical pipelines.