Anthropic spent millions buying millions of physical books. Then they had them shredded.
That’s not a metaphor. The AI company contracted a service to remove bindings, slice pages, scan every sheet—and dispose of the paper carcasses. The goal: train their models on text uncontaminated by AI-generated garbage or data poisoning. The method: destructive scanning, a legal loophole carved by a 2025 U.S. court ruling that equates destroying a physical copy with maintaining a “one-for-one” digital replica.
Code is truth. Intent is fiction. The outcome? A pipeline that turns cultural artifacts into fuel for large language models—and a business model that raises more questions than it answers.
The Context: Data Hungry, Legally Desperate
Every AI company faces the same problem: high-quality human-written text is finite. Web crawls are polluted with AI slop. Licensed datasets are expensive and narrow. So when a federal court ruled that converting a legitimately purchased physical book into a non-distributed digital copy—and then destroying the original—constituted fair use, a new market was born.

ISBNdb, a company that once sold book metadata, pivoted to become the middleman. They buy books by the pallet, filter by ISBN, publication year, or subject, then offer a “scan-and-destroy” service. Their marketing copy openly touts that pre-2022 physical books are less exposed to AI-generated text and modern data poisoning techniques. They even offer legally binding NDAs and verifiable destruction. Anthropic is their most famous client, confirmed as having paid millions for millions of volumes.
But this isn’t just a data acquisition strategy. It’s a systemic teardown of how we value physical knowledge.
The Core: A Systematic Teardown of Destructive Scanning
Let’s start with the technical mechanics. Buying a book, scanning it, then shredding it sounds straightforward—but the operational friction is enormous. Each book must be unbound, fed through high-speed scanners, OCR-processed, quality-checked, and stored. Then the paper must be disposed of in a way that satisfies the legal “destruction” requirement—shredding or incineration. The storage alone for millions of high-resolution PDFs adds petabytes of cloud costs. The scanning equipment and labor scale linearly with volume. I’ve audited data pipelines in DeFi that looked elegant on paper but collapsed under real-world throughput. This one is no different: the per-token cost of this method is almost certainly higher than licensing digital rights from publishers, yet companies like Anthropic choose it precisely because those rights are often unavailable or entangled in legal uncertainty.
The ledger keeps score. The court’s “one-for-one” reasoning is technically fragile. A digital copy can be replicated infinitely with a single command. The court assumed good-faith enforcement of non-distribution, but in practice, once the scan exists, the chain of custody is broken. A disgruntled employee, a backup leak, a cloud misconfiguration—any of these turns a legal dataset into a liability. The law treats intent as fiction, but the code treats replication as default.
Then there’s the cultural loss. The article notes that no specific titles of rare or unique books have been identified in destruction logs, but that’s precisely the problem: lack of evidence doesn’t mean lack of damage. ISBNdb’s own acknowledgment of “reputation concerns” around book burning underscores that the industry knows this looks bad. Libraries and archives now compete with AI companies for the same physical stock—but AI companies can outbid and then destroy. This isn’t a hypothetical; it’s a documented trend. The ethical calculus: sacrifice a unique physical artifact for a digital copy that will be used to train a model that may never even reference that book.
From a competition standpoint, this creates a moat—but a toxic one. Anthropic’s early move lets them lock up certain subject matters or eras of print. Later entrants will face either higher prices or depleted inventory. The result is a physical-world arms race with no finish line. Meanwhile, open-source models, which rely on publicly available datasets contaminated with AI-generated text, fall further behind. The disparity grows not because of algorithmic innovation, but because one side is willing to burn books.
The Contrarian: What the Bulls Got Right
To be fair, the proponents have a point. Physical books offer a signal-to-noise ratio that web text cannot match. No autogenerated spam, no SEO-driven keyword stuffing, no coordination attacks. For safety-critical applications—legal reasoning, medical diagnostics, historical analysis—the purity of source material matters. The court ruling provided legal clarity where none existed. ISBNdb’s service offers a clean, auditable chain of provenance. If an AI company can certify that every token in its training set came from a physically verified, human-authored book, that is a defensible quality claim.
But this logic ignores a crucial blind spot: the model doesn’t know it was trained only on old books. It will exhibit temporal bias, lack understanding of digital-native concepts, and inherit the prejudices embedded in pre-2022 publishing. The “clean data” argument is a technical illusion—purity in format does not equal purity in content.
Moreover, the legal foundation is shaky. The 2025 ruling is not binding nationwide, and the case against Anthropic over “pirated library copies” is still pending. A single appellate reversal could render millions of dollars in destroyed inventory a sunk cost—and the scanned copies a legal time bomb.
The Takeaway: Accountability Is Still Being Written
Minted nothing, promised everything. The AI industry is destroying physical books to obtain data that may not even be superior—just legally convenient. The ledger keeps score, but whose ledger? The culture that loses a rare edition cannot mint a replacement. The legal system that created this loophole may close it. The investors funding these book-burning campaigns should ask: what happens when the public learns that their AI was trained on ashes?
Gas fees don’t lie, but neither does the smell of burnt paper. In the long run, sustainable data strategies will not rely on irreversible destruction. They will license, collaborate, and digitize without pulping. Until then, every shredded spine is a promissory note with no collateral.