Anthropic’s Destructive Book Scanning Exposes AI Training’s Copyright and Cultural Stakes

Image: Theguardian
Main Takeaway
Anthropic digitized and discarded millions of print books for Claude training, a court case that has intensified debate over copyright, competition, and cultural preservation.
Jump to Key PointsSummary
Anthropic’s book-scanning project
Anthropic physically dismantled and digitized millions of print books to build training data for Claude, according to court documents and reporting on the company’s copyright litigation. The effort, known internally as Project Panama, involved cutting books from their bindings, scanning the pages, and discarding the originals after digitization.
The project emerged from a broader effort to obtain high-quality text created before generative AI became widespread. Books offered long-form prose, curated information, and varied language that Anthropic considered valuable for model training. The company hired Tom Turvey, a former Google Books partnerships executive, in February 2024 and assigned him a goal described in internal material as obtaining “all the books in the world.”
Why physical books mattered
Anthropic pursued print books because they provided a large body of pre-2022 writing that had not been generated or altered by AI systems. Training researchers have sought older text to reduce the risk of feeding models material produced by earlier models, while books remain unusually dense sources of sustained language and factual writing.
Destructive scanning also offered an operational shortcut. A high-speed process could separate pages, capture digital images, and remove the need to store or preserve bulky physical volumes. The approach raised a separate competitive concern: once a company controlled the resulting files, rival developers couldn't access the same collection easily. Anna’s Archive framed that dynamic as a privatization of cultural knowledge, while legal coverage focused on the process and its role in Anthropic’s fair-use case.
The legal dispute behind Project Panama
Project Panama became public through Bartz v. Anthropic PBC, a Northern California copyright case that examined how Anthropic obtained and used books for AI training. An internal memo described the initiative as an effort to “destructively scan all the books in the world” and advised employees to keep the codename private. Those details gave the scanning operation significance beyond a routine data-acquisition program.
The court’s ruling addressed whether Anthropic’s copying of books for model training qualified as fair use, while the treatment of purchased print copies created a distinct issue. Coverage of the decision described millions of dollars spent acquiring books and converting them into digital files. Anna’s Archive tied the episode to a reported $1.5 billion copyright settlement, while other accounts centered on the court record and the fair-use implications.
Cultural preservation and open access
The physical destruction of books has prompted a preservation campaign from Anna’s Archive, which is urging volunteers to scan and upload rare or endangered works. Its argument is that private digitization can leave society dependent on corporate servers for access to texts that once existed in public circulation.
That campaign adds a cultural dimension to a dispute usually described through copyright and model quality. A discarded book can be replaced when many copies survive, but rare editions, regional publications, and out-of-print works face a different risk. Public scanning projects also create their own legal questions, especially when uploading copyrighted books without permission. The conflict therefore pits corporate control, copyright enforcement, and preservation goals against one another.
What the episode means for AI training
Anthropic’s book operation shows how competition for high-quality training data has moved beyond web scraping and licensed datasets. Companies are pursuing curated collections, older text, specialist archives, and physical media as model developers search for material that remains useful and distinct from synthetic output.
The episode also puts procurement and data governance under scrutiny. Developers must track how training texts were acquired, whether copies were lawfully made, how long source material is retained, and whether the resulting datasets can be shared or audited. Courts will continue defining the boundary between transformative model training and unauthorized copying, while publishers and authors face pressure to negotiate clearer licensing terms.
What happens next
The immediate consequences will center on copyright litigation, licensing negotiations, and preservation efforts. Court findings about Anthropic’s process will inform arguments in other cases involving books, datasets, and AI model training, while publishers will examine whether access to curated text can command licensing revenue.
Anna’s Archive is pushing volunteers to scan books before private digitization removes them from circulation, but that response won't settle the underlying rights disputes. The central question is how society preserves written culture when AI companies can buy, digitize, and discard physical copies at industrial scale. Anthropic’s Project Panama has made that question concrete, attaching it to Claude’s development and to a legal record that now extends far beyond one company.
Key Points
Anthropic digitized and discarded millions of print books while building training data for Claude.
Project Panama recruited former Google Books executive Tom Turvey to expand Anthropic’s book acquisition operation.
Court documents brought Anthropic’s destructive scanning methods into a major copyright and fair-use dispute.
Anna’s Archive urged volunteers to preserve rare books through independent scanning and digital uploads.
The controversy links AI training-data competition with copyright licensing and cultural-heritage preservation.
Questions Answered
Anthropic cut apart and scanned millions of print books before discarding the originals for Claude training, according to court documents and reporting. The operation formed part of Project Panama, an internal book-digitization effort.
Anthropic’s Project Panama was a program to acquire and destructively scan large numbers of books for AI training data. An internal memo described the project as an effort to scan “all the books in the world.”
Anthropic used printed books to obtain high-quality, long-form text created before widespread generative AI output. Books offered curated language and factual content that researchers considered valuable for training Claude.
Anthropic’s book scanning raises copyright and fair-use questions about copying books to create AI training datasets. Bartz v. Anthropic examined the company’s acquisition, digitization, and use of those books.
Anna’s Archive says private destruction and digitization can place cultural knowledge under corporate control. It is urging volunteers to scan rare and endangered books for broader digital preservation.
Source Reliability
67% of sources are highly trusted · Avg reliability: 83
Go deeper with Organic Intel
Simple AI systems for your life, work, and business. Each one includes copyable prompts, guides, and downloadable resources.
Explore Systems