AI Companies Are Buying Millions of Old Books, Scanning Them for Training Data, Then Destroying the Originals

Image: 404media
Main Takeaway
AI developers including Anthropic are spending tens of millions on bulk book acquisitions, using destructive scanning to feed human-authored text into models while rare book dealers across Europe raise alarms.
Jump to Key PointsSummary
How the bulk book destruction pipeline works
AI companies have built an industrial pipeline for converting physical books into training data. The process, known as destructive scanning, involves purchasing books in bulk, removing their bindings with hydraulic cutting machines, scanning individual pages, and then discarding or shredding the originals. According to 404 Media, this method targets pre-2022 books specifically because they are free of AI-generated content, making them more valuable for training large language models.
Video footage obtained by News.com.au from ABTec Solutions, a Canadian digitization company, showed the process in action, with bindings being sliced off and pages fed through industrial-grade imaging equipment. The scale is staggering. Fortune reports that intermediaries like 2077AI are sending spreadsheets with over 3,000 ISBNs to antiquarian booksellers, requesting quotes and shipping estimates directed to China. Services facilitating these transactions advertise the ability to source up to one million volumes per order while keeping the end client anonymous.
Why pre-2022 books are the target
The preference for older, pre-2022 books is not about rarity, it is about data purity. As AI-generated content floods the internet, training datasets scraped from the web become increasingly contaminated with synthetic text. 404 News reports that ISBNdb, a company sourcing printed books for AI training, warns clients that the optics problem is real but emphasizes that old books provide a corpus of human-authored material untouched by AI slop.
This creates a perverse incentive structure. Books published before the generative AI boom represent a finite, irreplaceable resource of verified human writing. The Washington Post reported earlier this year that Anthropic, under the codename Project Panama, spent tens of millions of dollars acquiring millions of books to feed its demand for human-authored training material. The destruction of the originals after scanning means these physical artifacts are permanently removed from circulation.
Rare book dealers sound the alarm across Europe
Secondhand booksellers across Europe are raising concerns about this practice. NL Times reports that Dutch booksellers received requests from abroad, with Pieter de Vries of antiquarian bookstore De Vries & De Vries in Haarlem receiving an email from a woman identifying herself as Natalia from a company called 2077AI. The email included a spreadsheet with thousands of ISBNs and instructions to prepare quotes and coordinate shipping to China.
De Vries told Fortune he initially dismissed the message as spam or phishing. It was not until weeks later, when a Dutch journalist contacted him, that he realized the request was genuine and part of a broader pattern. Several other Dutch antiquarian booksellers received the same request. The concern among dealers is that obscure and often rare editions are being removed from circulation permanently, not just digitized.
The legal landscape and copyright implications
A federal judge has ruled that the digitization of purchased books, accompanied by destruction of the originals, falls within fair use under certain circumstances, according to court documents cited by Dallasexpress. This legal interpretation allows AI companies to argue that buying a physical book grants them the right to scan and discard it, since the text itself is not protected differently than the physical object once lawfully acquired.
The ruling does not settle all questions. Copyright lawsuits against AI developers continue to mount, and the anonymity of bulk purchasing through intermediaries complicates accountability. Yahoo Finance notes that a cottage industry has sprung up to supply AI companies with millions of physical books, with sourcing companies advertising their ability to locate hundreds of thousands of titles while keeping the end buyer's identity hidden. This opacity frustrates publishers and authors who want to know whose models are trained on their work.
A bookseller's perspective from inside the industry
Charlie Becker, a second-generation bookseller building an AI tool for used bookstores out of his family's store in Houston, offers a more nuanced view. Writing on his Substack, Becker describes pausing sales after receiving a single online order for 70 books, an order so unprecedented he needed to investigate. He questions whether an AI company is truly buying up all the used books, suggesting the current wave of concern may be overblown.
Becker acknowledges the broader stakes regardless of this specific moment. As he puts it, the Library of Alexandria is still burning. The metaphor captures the tension between the hunger for training data and the preservation of cultural heritage. Even if the current panic is exaggerated, the underlying dynamic is real and accelerating.
What this means for the future of books
The destruction of physical books for AI training data represents a collision between cultural preservation and technological acceleration. The books being destroyed are not just any books. They are often obscure, out of print, and rare editions that booksellers consider irreplaceable. Once scanned and shredded, the physical object is gone, even if the text lives on in a training dataset.
Russh frames this as arguably the most concerning disclosure in the artificial intelligence industry, taking place on our bookshelves. The process echoes the dystopian imagery of Fahrenheit 451, where books are not just read but burned. The difference is that the text survives, but only in a form inaccessible to human readers, locked inside a proprietary AI model's weights. The question that remains unanswered is whether the tradeoff is worth it, and who gets to decide.
What happens next
The practice is likely to continue as long as pre-2022 books remain the cleanest source of human-authored text and courts allow the destruction of lawfully purchased copies. The Dutch bookseller survey reported by NL Times suggests the phenomenon is spreading across Europe, not limited to North America. Booksellers are beginning to organize and share information about suspicious bulk orders.
Anthropic's Project Panama demonstrates that major AI companies are willing to invest tens of millions in this approach. The Washington Post report from earlier this year established the scale and seriousness of the effort. The question is whether public pressure or regulatory intervention will slow the practice. The optics problem that ISBNdb acknowledged to its clients is real, and it is not going away.
Key Points
Anthropic spent tens of millions on Project Panama to acquire and destructively scan millions of books for Claude's training data.
AI companies target pre-2022 books specifically because they contain no AI-generated text, making them clean training data.
Dutch bookseller Pieter de Vries received a request for 3,000 ISBNs from a company called 2077AI, which he initially dismissed as phishing.
Intermediaries like ISBNdb and 2077AI facilitate anonymous bulk book sourcing while warning clients about the optics problem.
A federal judge has ruled that destroying lawfully purchased books for digitization falls within fair use protections.
Questions Answered
AI companies are destroying old books after scanning them to create training data for large language models. The books are physically dismantled, their spines cut off, and pages scanned before the originals are discarded, because pre-2022 books contain no AI-generated content and provide clean human-authored text.
Yes, Anthropic spent tens of millions of dollars acquiring millions of books under the codename Project Panama, according to the Washington Post. The company used hydraulic cutting machines to remove bindings and scan pages before destroying the physical copies.
2077AI is a company that sources physical books for AI training data. It sent requests to Dutch antiquarian booksellers including Pieter de Vries, asking for quotes on over 3,000 ISBNs with instructions to ship the books to China.
A federal judge has ruled that digitizing lawfully purchased books and destroying the originals falls within fair use under certain circumstances. However, the broader copyright implications remain contested as lawsuits against AI developers continue.
Rare book dealers across Europe, particularly in the Netherlands, are raising concerns and sharing information about suspicious bulk orders. Some booksellers initially dismissed large requests as spam before realizing they were part of a systematic effort to source books for AI training data.
Source Reliability
44% of sources are trusted · Avg reliability: 70
Go deeper with Organic Intel
Simple AI systems for your life, work, and business. Each one includes copyable prompts, guides, and downloadable resources.
Explore Systems