Amazon, which started off selling books, is destroying rare texts to train AI Amazon, the e-commerce and cloud computing giant, is purchasing and physically destroying rare, out-of-print books to digitize their contents for training large language models, according to TechCrunch. The company has established a dedicated scanning facility to process these texts, aiming to access a vast, uncontaminated corpus of pre-2022 human writing to prevent model collapse and improve output quality. This move gives Amazon a proprietary, non-reproducible dataset that competitors relying on public-domain or licensed data cannot match. TechCrunch https://techcrunch.com/2026/08/17/amazon-once-an-online-bookseller-is-destroying-rare-books-to-train-ai-models/ Amazon, which started off selling books, is destroying rare texts to train AI Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated. Amazon is now physically destroying rare, out-of-print books to feed their LLM training pipelines. This unlocks a vast, uncontaminated corpus of pre-2022 human text that can prevent model collapse and improve output quality, but it also means any team relying on public-domain or licensed data is now competing with a proprietary, non-reproducible dataset that only Amazon can build. Amazon is systematically purchasing and destroying rare, out-of-print books to digitize them for LLM training at a dedicated physical scanning facility. This aggressive offline ingestion highlights the critical shortage of clean, pre-2022 human text, signaling to production teams that the high-quality, non-synthetic data required to prevent model collapse is increasingly gatekept behind proprietary physical pipelines.