New investigations reveal that large-scale AI training may involve the destruction of physical books. Learn the ethical and legal implications of this practice.

Recent investigations confirm that some tech firms are physically destroying rare books to expedite AI training. At Groundwork, we emphasize that this practice raises critical questions about copyright, cultural preservation, and corporate transparency. Stakeholders should demand clearer disclosure standards for AI training datasets.
Based on reporting by Ars Technica. Research, structure, and fact-checking by Groundwork.
“This scenario underscores the tension between rapid technological scaling and the preservation of physical intellectual property. Groundwork's analysis suggests that without clear legal frameworks, companies will continue to prioritize data throughput over the long-term integrity of our cultural record.”
Data scraping for artificial intelligence training refers to the systematic collection and digitization of physical or digital content to feed large language models (LLMs). Recent investigations, including reports from 404 Media and Ars Technica, have highlighted the physical destruction of rare and copyrighted books within corporate facilities to facilitate high-speed scanning and ingestion into AI training pipelines. This practice raises significant questions regarding intellectual property rights, the preservation of cultural heritage, and the transparency of proprietary model development.
AI model training requires massive datasets to identify patterns, linguistic nuances, and factual information. While much of this data is scraped from the open web, tech companies increasingly turn to physical media to access high-quality, long-form content that may not be available in digital repositories. At Groundwork, our analysis shows that the process often involves the 'guillotine-style' removal of book spines to enable high-speed automatic document feeders (ADFs). By converting physical pages into machine-readable text via optical character recognition (OCR), companies can ingest vast amounts of copyrighted literature into their training sets without the traditional limitations of digital licensing or public domain access.
Evidence surfaced through the use of tracking technology—specifically, an AirTag planted in a rare book—confirmed that bulk orders of literature were being routed to specific AI training facilities. According to 404 Media, the tracking data led investigators to a warehouse in Las Vegas, identified as 'VGT3,' where books were systematically deconstructed. The presence of internal branding depicting a dinosaur devouring a book underscores a corporate culture that prioritizes the rapid acquisition of training data over the preservation of the physical artifacts. This finding serves as a primary data point in the ongoing debate over whether AI development constitutes 'fair use' under current copyright law.
Copyright law provides authors and publishers with the exclusive right to reproduce their work. When tech firms purchase books and physically destroy them to digitize the content, they are essentially creating a derivative digital copy. While companies often argue that the resulting training data is transformative, legal experts remain divided. The destruction of unique or rare copies adds a layer of ethical complexity, as it effectively removes the original artifact from the secondary market, potentially impacting the value of remaining copies and the provenance of the work. As of 2026, several high-profile lawsuits are challenging whether the ingestion of copyrighted works into LLMs without compensation violates existing intellectual property statutes.
Transparency in AI development is currently minimal, as most companies treat their training datasets as protected trade secrets. When consumers and authors are unaware of how their data—or the cultural artifacts they value—is being utilized, trust in the AI ecosystem erodes. At Groundwork, we emphasize that the lack of disclosure regarding source material creates a 'black box' effect. If a company refuses to confirm whether specific training methodologies involve the destruction of physical media, it complicates the ability of regulators and the public to hold these organizations accountable for their supply chain practices.
Alternatives to the destruction of physical media exist, though they are often more expensive or time-consuming for corporations. These include:
For authors, publishers, and consumers concerned about the future of intellectual property, the path forward involves rigorous advocacy for transparency. This includes supporting legislative efforts that require companies to disclose the sources of their training data and demanding that AI firms adhere to ethical acquisition standards. By tracking the flow of physical goods and maintaining pressure on tech conglomerates, the public can demand a higher standard of accountability in the pursuit of artificial intelligence advancement. Decisions regarding the future of knowledge and culture should not be made in the dark; they require a balanced approach that respects both innovation and the sanctity of the written word.
Sofia Reyes (2026). The ethics and mechanics of data harvesting in AI model training. Groundwork. Retrieved from https://gworky.com/article/amazon-book-destruction-ai-training-investigation
Evidence-based verification conducted by the Groundwork Research Desk
Groundwork enforces a strict, independent verification standard. Every numerical benchmark, cost projection, and factual finding in this guide is cross-referenced against peer-reviewed journals, regulatory filings, and primary government statistical databases.
The legality of this practice is currently being litigated in multiple courts. While the purchase of a book grants ownership of the physical object, it does not necessarily grant the right to reproduce or distribute the contents, even for the purpose of training an AI model.
Currently, there is no public registry that allows authors or owners to verify if their specific books have been ingested into an AI training set. Companies treat their datasets as proprietary trade secrets, making it difficult for individuals to confirm the origin of the data used in their models.
AI firms often require physical books because they contain high-quality, long-form text that has not been digitized or is locked behind digital rights management (DRM) systems. Physical books provide a reliable source of clean, structured text that is ideal for training complex language models.
The primary concern is the permanent loss of cultural artifacts. When rare books are destroyed to create machine-readable data, the original version of the work is removed from the market, which can devalue the history of the work and limit access for future researchers and collectors.
Smart Home & Digital Privacy Analyst
Smart home and digital privacy analyst focused on data ownership, device security, and power efficiency of AI utilities and gadgets.
This guide underwent secondary data verification to confirm primary source integrity, calculation formulas, and regulatory compliance before publication.

Perplexity's partnership with Airtel provides a case study on AI growth experiments. We analyze the effectiveness of subsidized scaling and user retention.
FLOPs are a common but flawed way to measure AI efficiency. Learn why they fail to predict real-world performance and how to use empirical benchmarks instead.
Learn how using KL divergence for principled gating in multi-agent reinforcement learning improves coordination stability and reduces communication noise.