Skip to content

AI labs are buying books by the pallet, and destroying them to scan them.

A broker sells labs books printed before 2022, guaranteed to contain no generated text. High-speed scanning requires cutting off the spine. Rare editions disappear in the process.

Advertisement
The essentials in 30 seconds ⚡
AI labs are buying printed books from before 2022, in batches ranging from a thousand to a million copies, via intermediaries bound by confidentiality agreements. These works are then scanned at high speed, which usually involves cutting off the spine to feed the pages through the machine. The physical book does not survive the process. The reason for this rush comes down to two words: clean data.

There are topics where outrage comes too quickly and gets in the way of understanding. This is one of them. Let's take the time to explain why this is happening, before saying what problems it raises.

Why pre-2022 books suddenly hold gold

The main reason, we described it in our article on model collapse. A model trained on content produced by other models gradually degrades: it loses diversity, repeats errors, and hallucinates more. It's the photocopy of a photocopy.

Yet the web has been massively flooded with machine-generated text since 2022. For a lab, distinguishing a page written by a human from one produced by a machine has become difficult and costly. Hence the appeal of a source whose dating guarantees its origin: a book printed before the arrival of consumer-grade generators cannot contain synthetic text.

The book also has three advantages the web lacks. It has been edited and proofread, so its linguistic quality is superior to that of an average page. It is structured, with an argument unfolding over hundreds of pages, which is valuable for learning to reason at length. And many works have never been digitised, making them fresh material in a sector where the easy data has already been tapped.

The second reason, which few people know about 🧪
Authors have started fighting back through data poisoning: inserting into their texts elements invisible to a human reader but disruptive to a model during training. The broker selling these books specifically highlights that a very small number of carefully crafted documents can be enough to install unwanted behaviour in a corpus that is nonetheless gigantic. A book printed before these tools existed escapes the problem by construction. The selling point is therefore not just the cleanliness of the text, it's also its traceability: buying the paper, keeping the invoices, and having a chain of custody that lawyers can defend.

What actually happens

The reported mechanism is industrial. Orders go through intermediaries, under strict confidentiality agreements, for volumes ranging from a thousand to a million copies. The books arrive on pallets, sometimes over a thousand per pallet.

Then comes digitisation. Scanning a book without damaging it exists, but it's slow and costly: each page has to be handled. The fast method involves cutting off the spine to get a stack of loose sheets, which goes through an automatic scanner like any document. It's far quicker and far cheaper. The book, meanwhile, is destroyed.

The phenomenon has been observed in several European countries, including the Netherlands, Germany, Switzerland and Spain, where second-hand booksellers report sudden requests for bulk purchases where the only criterion is volume, with no regard for the value of the works.

What this really threatens

Let's be precise, because not everything is equally serious.

Destroying a common copy is not a tragedy. A bestselling novel printed in hundreds of thousands of copies does not disappear because a thousand are scanned. It's an object, not a work.

The problem lies elsewhere: rare books. When the purchase criterion is volume rather than content, out-of-print editions, limited print runs, and technical or regional works never digitised go into the lot. And for those, the destruction can be permanent. The open question, which no one can answer, is how many editions that existed nowhere else have already gone under the blade of an industrial scanner before anyone could digitise them properly.

The third problem is economic. This demand creates a financial incentive for destruction. A second-hand bookseller offered to offload an entire stock by weight has no obvious reason to refuse. The market for old books has never had to contend with a buyer whose interest lies in making the merchandise disappear.

And the law in all this?

This is the most counter-intuitive point. Buying a book and scanning it for your own use is not the same as downloading a pirated copy. A widely discussed US court ruling has in fact distinguished between these two situations: training from legally acquired works was deemed transformative, while using illegally obtained copies cost one company a settlement of several hundred million dollars with authors.

It's precisely for this reason that the second-hand market has become attractive: it offers a legally defensible path where downloading is not. The destruction of the book is, paradoxically, a consequence of the concern for legal compliance.

That doesn't exhaust the debate, though. The proceedings we are following, notably the one pitting publishers against Google, concern similar questions and have not yet produced a stable doctrine.

What this story really says

Beyond the striking image of shredded books, this episode reveals a fundamental shift in the AI economy. For years, data was abundant and free: you just had to scrape the web. That period is closing, for three simultaneous reasons. Clean text is becoming scarce because the web is getting polluted. Rights holders are fighting back, through the courts and through technology. And what remains is expensive to obtain.

The consequence is that data is becoming a strategic resource again, contested, with owners. It's the same movement we observed regarding licensing deals signed with press or music publishers. The difference here is that the resource is physical, finite, and sometimes irreplaceable.

There's an irony hard to ignore to close. These systems were built by absorbing human writing. Now they're exhausting the accessible reserve, to the point where you have to go get the paper from back rooms and destroy it to read it. You could see it as an easy symbol. You could also see it as a useful reminder: whatever these machines know, they owe to us, and that reserve was not infinite.

Advertisement