AI companies keep destroying old books. Here's why.
Anthropic, the company behind Claude, reportedly purchased and physically destroyed millions of books as part of an effort to obtain high quality text for training its artificial intelligence models. Court documents revealed an internal initiative known as “Project Panama,” which aimed to build an enormous digital library for AI research and training.
Beginning in 2024, Anthropic spent millions of dollars buying physical books, many of them used. Contractors removed the bindings, cut the pages into suitable dimensions and scanned them using high speed equipment. After digitisation, the physical copies were discarded rather than preserved.
The strategy reflected the AI industry’s demand for professionally written material. Books are particularly valuable for training large language models because they contain edited, structured and generally higher quality writing than much of the material available freely on the internet. Anthropic executives believed such material could improve Claude’s writing and reasoning capabilities.
The legal background is important. Anthropic had previously obtained millions of digital books from pirate libraries including LibGen. Authors subsequently sued the company for copyright infringement. In 2025, a federal judge distinguished between these pirated copies and physical books that Anthropic had legally purchased and subsequently scanned.
Judge William Alsup ruled that Anthropic’s process of purchasing physical books, converting them into digital copies and destroying the originals qualified as fair use. The ruling effectively treated the process as changing the format of a legally acquired copy rather than creating additional copies for distribution. However, the company’s acquisition of books through pirate libraries raised separate copyright issues.
Anthropic later agreed to a $1.5 billion settlement with authors over allegations concerning pirated books used in developing its AI systems. The settlement was expected to cover roughly 500,000 works, illustrating the enormous legal and financial stakes surrounding AI training data.
The revelations have also created broader concerns about preservation. Recent reports indicate that booksellers are seeing unusually large orders for obscure and older books, leading to fears that AI companies or intermediaries may be acquiring additional physical books for scanning and destruction. However, there is currently no definitive evidence connecting all of these recent purchases directly to Anthropic or other specific AI companies.
Overall, the controversy illustrates the extraordinary lengths AI companies are willing to go to secure high quality training data, while raising difficult questions about copyright, compensation for authors, transparency and the preservation of physical books.





