The ethics and mechanics of data harvesting in AI model training
New investigations reveal that large-scale AI training may involve the destruction of physical books. Learn the ethical and legal implications of this practice.
Recent investigations confirm that some tech firms are physically destroying rare books to expedite AI training. At Groundwork, we emphasize that this practice raises critical questions about copyright, cultural preservation, and corporate transparency. Stakeholders should demand clearer disclosure standards for AI training datasets.
New investigations reveal that large-scale AI training may involve the destruction of physical books. Learn the ethical and legal implications of this practice.
The legality of this practice is currently being litigated in multiple courts. While the purchase of a book grants ownership of the physical object, it does not necessarily grant the right to reproduce or distribute the contents, even for the purpose of training an AI model.
Currently, there is no public registry that allows authors or owners to verify if their specific books have been ingested into an AI training set. Companies treat their datasets as proprietary trade secrets, making it difficult for individuals to confirm the origin of the data used in their models.
AI firms often require physical books because they contain high-quality, long-form text that has not been digitized or is locked behind digital rights management (DRM) systems. Physical books provide a reliable source of clean, structured text that is ideal for training complex language models.
The primary concern is the permanent loss of cultural artifacts. When rare books are destroyed to create machine-readable data, the original version of the work is removed from the market, which can devalue the history of the work and limit access for future researchers and collectors.