Amazon and other tech firms are scanning rare books for AI training, often destroying the physical copies. Learn the implications for data quality and ethics.

AI companies are physically destroying rare books to create 'clean' training data for LLMs, aiming to avoid the quality degradation caused by AI-generated content. While this ensures high-quality training sets, it permanently destroys unique cultural artifacts, raising significant ethical and legal questions about data sourcing.
Based on reporting by TechCrunch Enterprise & AI. Research, structure, and fact-checking by Groundwork.
“This practice illustrates the extreme measures tech firms are taking to combat model collapse, a known technical failure in AI development. The destruction of physical archives is an empirical indicator that high-quality, human-generated data is becoming a scarce, non-renewable resource in the AI era.”
AI training data refers to the massive datasets, including books, articles, and internet content, used to teach Large Language Models (LLMs) to recognize patterns and generate human-like text. As companies exhaust easily accessible digital sources, they have increasingly turned to physical archives, including rare and out-of-print books, to feed their models. Recent investigations, including reports by 404 Media, have revealed that major technology companies are purchasing rare texts, removing their bindings, and scanning them systematically to convert physical pages into digital training data.
At Groundwork, our analysis shows that this practice represents a significant shift in how AI developers source information. When digital repositories are depleted, the physical world becomes the final frontier for high-quality data. By targeting older, rare books—specifically those published before 2022—developers aim to avoid 'model collapse,' a phenomenon where AI models degrade after training on low-quality, AI-generated content.
AI companies target rare books because they provide high-quality, human-authored data that predates the proliferation of AI-generated content. As LLMs become more prevalent, the internet is increasingly saturated with synthetic text, which can introduce recursive errors and biases into new models. By sourcing books published before 2022, developers ensure their training sets contain authentic human linguistic nuances and historical accuracy that synthetic data cannot replicate.
According to research on model collapse, training models on AI-generated content leads to a rapid decline in output quality, often resulting in 'gibberish' or loss of nuance (Journal of Machine Learning Research). Rare books represent a 'clean' data source, free from the feedback loops that plague modern web-scraped datasets. For developers, the value of this data outweighs the destruction of the physical artifact.
The process of digitizing rare books for AI training involves high-speed industrial scanning, which necessitates the destruction of the book’s physical structure. The workflow typically follows these steps:
This industrial-scale digitization transforms a limited-edition object into a commodity data asset. Once the physical book is scanned, it is rarely, if ever, preserved in its original form.
The destruction of rare texts for AI training raises profound questions regarding intellectual property, cultural preservation, and the definition of a 'fair' use of information. While AI companies argue that they are simply 'improving products' through commercial data acquisition, critics point to the loss of irreplaceable cultural artifacts. When a book is destroyed to be digitized, the unique historical context of that specific copy—annotations, marginalia, or provenance—is permanently erased.
At Groundwork, we find that the tension lies between the democratization of information and the destruction of physical heritage. While digitizing a text can make its content accessible to millions, the destruction of the original copy removes the ability for researchers, historians, and collectors to verify the material or examine the artifact itself. This creates a trade-off where the content is preserved, but the medium is sacrificed.
AI companies could theoretically rely on existing digital archives like the Internet Archive or Project Gutenberg, but these sources often face legal and technical limitations. Many older books are protected by copyright, and their digital availability is often restricted by licensing agreements or legal injunctions. By purchasing physical copies and scanning them, companies argue they are exercising ownership rights over the physical items they have legally acquired.
However, the legal landscape is shifting. Several ongoing lawsuits, including those involving authors and publishers against AI developers, challenge the premise that purchasing a book confers the right to use its contents for machine learning training. Until these legal frameworks are solidified, companies are continuing to operate in a gray area, prioritizing the acquisition of 'clean' data over the preservation of physical inventories.
The reliance on rare books for AI training signals an impending 'data drought' for high-quality, human-authored information. As the pool of un-scanned, pre-AI text shrinks, companies may be forced to rely on synthetic data, incentivized by the need to maintain model performance despite the risks of degradation.
For the consumer, this trend highlights the importance of supporting physical libraries and independent archives that prioritize the longevity of the book as an object. If your interest lies in the long-term preservation of knowledge, understand that the commercialization of data has turned even the most obscure, out-of-print books into high-value assets for technology conglomerates. Moving forward, expect to see more scrutiny regarding the origins of training datasets and increased pressure on companies to disclose whether their 'commercial channels' involve the destruction of cultural property.
Sofia Reyes (2026). The ethics and mechanics of using rare books for AI training. Groundwork. Retrieved from https://gworky.com/article/amazon-rare-book-destruction-ai-training
Evidence-based verification conducted by the Groundwork Research Desk
Groundwork enforces a strict, independent verification standard. Every numerical benchmark, cost projection, and factual finding in this guide is cross-referenced against peer-reviewed journals, regulatory filings, and primary government statistical databases.
AI models require rare books because they provide 'clean' data written before the rise of AI-generated content. Training on existing internet data risks 'model collapse,' where the AI learns from its own low-quality output, leading to degradation. Rare, pre-2022 texts ensure the model learns from authentic human language.
The legality is currently debated in court, though companies argue that purchasing books through commercial channels grants them the right to use the contents. However, lawsuits from authors and publishers are challenging whether this constitutes fair use or copyright infringement under current intellectual property laws.
In most industrial scanning processes, the books are destroyed. The spines are cut off to feed the pages into high-speed scanners, making the original physical volume unusable and effectively discarded once the digital capture is complete.
While digital libraries exist, they are often subject to strict licensing and copyright restrictions. AI companies prefer to own the physical copies to avoid the legal complexities of digital licensing, allowing them to process the text directly for their proprietary training pipelines.
Smart Home & Digital Privacy Analyst
Smart home and digital privacy analyst focused on data ownership, device security, and power efficiency of AI utilities and gadgets.
This guide underwent secondary data verification to confirm primary source integrity, calculation formulas, and regulatory compliance before publication.

Cargo theft targeting AI hardware has turned violent. Groundwork analyzes how criminal syndicates are bypassing security and how logistics firms can respond.

A forensic study of a skeleton at Stirling Castle reveals the only known trebuchet casualty in history, offering a rare look at medieval siege trauma.

Nvidia’s $21 billion stake in SpaceX represents a major strategic move to secure its role in satellite AI and compute infrastructure. Learn the implications.