To train AI we are literally tearing apart millions of books: the dark side of the data race

Volumes purchased in bulk, stripped of binding with industrial machinery, scanned page after page and finally sent to waste. It’s not the plot of a new one Fahrenheit 451 (1953, pure dystopia by the American writer Ray Bradbury)but one of the methods used to obtain data with which to train artificial intelligence.

The story was brought to light last January Washington Postbased on thousands of pages of court documents related to Anthropic, the company developing the language model Claude.

All documents that describe an internal project with an almost cinematic name: Project Panama, started in 2024 with the aim of procuring enormous quantities of books to be transformed into data for training artificial intelligence models.

A subsequent investigation by 404 Media shows how paper books, especially those published before the explosion of generative AI, are becoming an increasingly attractive commodity for companies looking for training material written by humans. A company that operates in the sector, ISBNdb, explicitly presents books as a particularly valuable source of data because they are composed of edited, structured human knowledge and, above all, free of the so-called “AI slop”, the mass of synthetic content that now invades the web.

The paradox is served: artificial intelligence needs books because the internet is increasingly full of content produced by artificial intelligence itself.

What is Project Panama

Project Panama was born in early 2024. According to court documents reviewed by the Washington PostAnthropic reportedly spent tens of millions of dollars in about a year to procure and digitize millions of books. An internal document explicitly described the project as an attempt to destructively scan a huge amount of volumes and stated that the company preferred not to make the operation public.

The technique used is called destructive scanningdestructive scanning: in practice, the binding is removed, the spine cut and the pages, now separated, quickly passed through industrial scanners. In the end, the digital file remains. The physical volume, however, can no longer be recomposed and is sent for disposal or recycling.

The same federal judge who handled the case describes in his decision books regularly purchased by Anthropic from which the bindings were removed, then scanned page by page and preserved as searchable digital files.

In other words, to preserve the text you destroy the object that contains it.

Why books?

Because a book is much more interesting to an AI than millions of mediocre web pages. In internal Anthropic documents cited by Washington Postas early as 2023, one of the founders hypothesized that training models on books could teach them to “write well,” instead of absorbing the low-quality language found online.

For years the internet has been considered the inexhaustible mine from which to extract texts, images and knowledge. But the arrival of generative AI systems has also progressively changed what is found online.

Automatically produced articles, reviews, commercial descriptions, translations, summaries, SEO posts, synthetic images: a growing share of online content no longer originates directly from human beings. And so what until yesterday seemed antiquated – an old printed book forgotten on a shelf – suddenly becomes a very precious resource.

404 Avg it tells the story of the emergence of this market. ISBNdb offers companies services for the acquisition of large quantities of books and underlines the value of editorially curated texts which certainly precede the proliferation of content generated by AI.

Now, if a new generation of models is increasingly trained on content produced by previous generations, errors, simplifications and distortions can propagate into the material used for new training. It’s one of the reasons why high-quality, clearly human-produced datasets are becoming so valuable. And this is why books, especially those published before the boom of ChatGPT and other generative systems, become attractive.

But is it legal to buy a book, scan it and destroy it? In the United States, at least in the Anthropic case, the answer given by a court was partly yes.

In Bartz vs. Anthropic, federal judge William Alsup ruled in June 2025 that using the works to train Claude and his predecessors could fall under the fair usefinding the training highly transformative. He also considered fair use also the digitization of paper books purchased legally, because Anthropic had replaced each physical copy owned with an internal digital copy, without placing further copies on the market.

But buying a book and digitizing it is not the same as illegally downloading a copy and the judge has in fact ruled out that the creation of a permanent library through pirated copies could be justified with the fair use. Anthropic had also acquired millions of files from so-called shadow librariesonline archives containing protected works disseminated without authorization.

The piracy controversy then led to a $1.5 billion settlement covering more than 480,000 works, an agreement that received final court approval in the summer of 2026. A distinction that is anything but secondary: there is not, even in the United States, a general pass to take any book and use it as one wishes.

And in Europe?

The European Union does not have the general equivalent of fair use American. The European copyright directive instead provides a specific exception for text and data mining, i.e. the automated extraction of information from large quantities of works to which there is legitimate access.

Article 4 of Directive 2019/790 allows certain reproductions and extractions for text and data mining, but also provides that rights holders may expressly reserve the use of their works, the so-called opt-out.

Added to this today is the AI ​​Act: from 2 August 2025, providers of general purpose AI models placed on the European market must adopt a policy aimed at respecting Union copyright, including the recognition of reservations expressed by rights holders, and publish a sufficiently detailed summary of the content used for training. Furthermore, from 2 August 2026 the related enforcement activity foreseen for these obligations entered the full application phase.

Now, however, the fundamental issue remains understanding how the rights of authors must be concretely respected in a technology that requires gigantic quantities of texts and data. For centuries we have built libraries so that books could be stored and read by people.

Now companies with enormous capital are building more libraries: invisible, digital and designed first and foremost to be read by machines. As for cultural heritage, how long will we be willing to transform it into private raw material in order to build the artificial intelligence of the future?