Massimo@Rainmaker1973
AI companies are buying large volumes of used, rare, and out-of-print books, scanning them for training data, and in many cases destroying the physical copies afterward.
This practice has accelerated as labs seek high-quality, pre-2022 human-written text free of AI-generated content that now saturates much of the open web.
The most documented example is Anthropic’s internal effort known as Project Panama. Court filings from copyright litigation revealed that the company acquired millions of physical books from used-book sellers and wholesalers such as Better World Books. Workers used hydraulic cutting machines to slice off the spines, fed the loose pages into industrial high-speed scanners, and then discarded or pulped the remains. An internal planning document described the goal as an effort to destructively scan books at massive scale. A federal judge later ruled that purchasing the books, creating internal digital copies this way, and destroying the originals constituted transformative fair use under U.S. copyright law, relying in part on the first-sale doctrine.
Similar activity is now widespread. Specialized brokers, including ISBNdb, openly market bulk acquisition services to AI labs, offering orders ranging from thousands to as many as one million titles. They emphasize older printed books as “structurally clean” training data.
Booksellers in the United States and Europe have reported sudden surges in bulk orders for mixed, random selections that include rare and out-of-print volumes. Some of these titles may represent among the last readily available physical copies. While a substantial portion of the books involved would otherwise have been remaindered, recycled, or sent to landfills, the inclusion of uncommon editions has drawn criticism from archivists, collectors, and rare-book dealers who note that gentler, non-destructive scanning methods exist yet are slower and more expensive.
The process remains largely quiet because the intermediary services promise anonymity to their AI clients. Public attention has grown mainly through litigation disclosures and reporting by outlets that examined the court records and interviewed booksellers.
The practice highlights a practical tension in large-scale AI development: the demand for vast quantities of reliable human text collides with the finite physical supply of older printed works and the cultural value of preserving rare copies.