AI ONLINE6 September 2026
The AI News Desk

RelayON THE WIRE

The whole field of AI — read, checked, and explained.
Policy & Safety

AI Companies Are Shredding Books to Train Their Models. A Shadow Library Is Racing to Scan Them First.

Anthropic's court-documented 'Project Panama' set out to 'destructively scan all the books in the world.' It's legal and cheap — and, for rare titles, a preservation problem. Anna's Archive is pointing the same scanners the other way.

Morgan ValeBy Morgan ValeSenior Desk Writer
21 August 2026
Listen to this postread by Relay

A Dutch second-hand bookseller received an order this summer for around 3,000 books — a spreadsheet of some three thousand separate titles — and assumed it was spam or a phishing scam. Who buys three thousand used books at once? The answer, increasingly, is an AI company — and the books are not being read. They are being sliced out of their bindings, run through a high-speed scanner, and shredded.

The practice has a name inside at least one lab. Court filings unsealed in the copyright litigation against Anthropic describe "Project Panama," an effort that began in early 2024 and was summarised in an internal planning document, in about as blunt a sentence as the industry has produced: the goal was "to destructively scan all the books in the world." One six-month contract alone planned for between 500,000 and two million books.

Why a chatbot wants a paper book

The demand is a side-effect of a problem the labs created for themselves. The open web, the traditional feedstock for training data, is now heavily contaminated with text written by earlier AI models — "slop," in the trade — which makes it worse than useless for teaching a new model to write like a person. Books published before 2022 are a clean reservoir of human-authored, professionally edited prose, and millions of out-of-print titles exist nowhere in digital form at all.

Buying the physical copies and scanning them also sidesteps a specific legal landmine. Anthropic learned the cost of the alternative the hard way: it agreed, in 2025, to a $1.5 billion settlement over a separate repository of roughly seven million pirated books. Purchasing used copies and digitising them is treated very differently by the courts — a federal judge ruled that scanning bought-and-paid-for books, even when the paper original is then destroyed, is transformative fair use, and the first-sale doctrine has long let whoever owns a physical book do what they like with it, including bin it.

So the model that emerges is legally clean and, to the companies, obviously worth it. The books cost a few dollars each; the training signal is scarce.

Is this "book burning"?

The story has travelled under headlines about AI firms "burning" books, and that framing is worth handling carefully. Most of what is being destroyed is ordinary second-hand stock. Anthropic's Panama orders ran to bulk academic and used paperbacks that exist in thousands of identical copies, and the company has said on the record that none of its data-buying programmes buy and destroy "rare" or "antiquarian" books. Calling that arson is overheated.

The sharper worry is narrower, and it points somewhere else. Investigative reporting by 404 Media hid a tracker in a shipment of rare books — close to a thousand of them — and followed it not to Anthropic but to an Amazon scanning facility in Las Vegas. When a title survives in only a handful of copies and one of them is guillotined for its data, the scan may capture the words but the object is gone, and with it anything a future scholar might have wanted that a flatbed cannot record: marginalia, bindings, editions, provenance. The scan is not a preservation copy. It is a byproduct of consumption.

The mirror-image campaign

Which is what makes the counter-move interesting. Anna's Archive — the shadow library whose viral post this week put the story back in front of a large audience — is urging a race to scan rare and fragile books for preservation, before the same titles are bought up and destroyed for training runs. It is the identical technology pointed in the opposite direction: one side scans to consume the original, the other to save it.

That symmetry is the honest shape of this story, and it is more uncomfortable than a simple villain. The copyright fight over AI training data has mostly been argued over who owns the words. This is a quieter question about the paper — what we lose when the cheapest way to feed a model is to disassemble the one physical thing that held a text, and whether anyone is keeping a copy for reasons that have nothing to do with a training run.

Tune your feed
Like to get more stories like this in your For You feed — dislike for fewer.
Sources
Morgan Vale — Senior Desk Writer. Morgan writes the clear, no-jargon explainers — the pieces that turn a dense launch or paper into something you can actually use. Spot something wrong? Tell me and I'll correct it in public.
Got a question about this?

Ask Relay — he reads every question himself and replies personally by email.

Ask Relay →