To legally train artificial intelligence systems, technology companies are buying millions of physical books—just to slice off their bindings, scan the pages, and shred the originals. A June 2025 federal ruling did not stop this practice. In some ways, it quietly codified it.
🔑 In brief
- The Bartz v. Anthropic case validated format shifting of lawfully purchased books as fair use.
- Vendors such as ISBNdb now sell up to one million physical books per engagement to AI developers.
- Print books published before 2022 became strategic feedstock because they avoid synthetic-text pollution.
- The one-copy-replaces-one-copy logic incentivizes destruction of the physical originals.
- No federal legislation has yet created a mandatory licensing scheme for AI training on books.
A landmark ruling that legalizes destruction
The case at the center of this storm is Bartz v. Anthropic, a class action brought by authors Andrea Bartz, Charles Graeber, and Kirk Wallace Johnson against AI company Anthropic PBC, creator of Claude. The plaintiffs accused the company of copying their works without authorization—both via pirated downloads from LibGen and Pirate Library Mirror, and by purchasing physical books that were then cut, scanned, and destroyed—to build an internal digital library used for training. In June 2025, Judge William Alsup of the U.S. District Court for the Northern District of California issued a ruling that reshaped the legal landscape for AI training data.
The court reached two distinct conclusions. First, it held that using copyrighted books to train large language models constituted fair use because the process extracts statistical relationships between text fragments rather than reproducing or distributing the original works. The judge described the use as « spectacularly transformative »—a phrase quickly cited across the tech legal literature. Second, and more consequentially for the book trade, the court ruled that converting lawfully purchased print books into digital library copies was also fair use, on the grounds that each PDF replaced the purchased copy without increasing the total number of copies held. One legal copy simply took the place of another. The paper original had to go.

Project Panama: Anthropic’s scanning factory
This second holding transforms book destruction into a legal elegance. Anthropic spent tens of millions of dollars acquiring millions of print books, often in used condition. Specialized service providers then stripped the pages from their bindings, cut them to size, and scanned them—discarding the paper. Each physical book yielded a PDF containing scanned page images with machine-readable text. Anthropic’s own internal documents, revealed during discovery, referred to this program as Project Panama. Tom Turvey, former head of partnerships for Google’s book-scanning project, was hired in February 2024 specifically to develop a legally defensible route to acquiring « all the books in the world » while avoiding « legal, practice, business slog. »
« Every purchased print copy was copied in order to save storage space and to enable searchability as a digital copy. The print original was destroyed. One replaced the other. »
Judge William Alsup, Bartz v. Anthropic
The one-copy-replaces-one-copy logic creates a perverse incentive structure. Keeping both the book and its scan would present a different copy-count fact pattern. Destroying the physical original, by contrast, preserves the legal premise that one owned copy has been exchanged for one digital copy. Preservation becomes not just unnecessary but potentially legally inconvenient. A company seeking maximum legal protection receives a clear directive: buy, scan, then dispose of the evidence.
ISBNdb: the commercial infrastructure of destruction
The ruling immediately opened a market. ISBNdb, a company that describes itself as the world’s largest book database, began marketing bulk physical book acquisition services directly to AI developers. The pitch is blunt: « The world’s best AI training data is sitting on a shelf. » Books, the company explains, are « dense, edited, authoritative »—characteristics increasingly rare in online text polluted by generative AI output.
| Characteristic | Pre-2022 print books | Open web today |
|---|---|---|
| Synthetic text risk | Near zero | High and rising |
| Exposure to data poisoning | None (print) | High (Nightshade, etc.) |
| Editorial quality | Filtered by publishers | Heterogeneous |
| Procurement traceability | Invoice, ISBN | Often impossible |
The service allows orders of between 1,000 and one million physical books per engagement, with a catalog that includes non-digitized, rare, and out-of-print titles. A sourcing article on the site promotes a legally binding nondisclosure agreement for each engagement and describes destructive scanning followed by verifiable destruction or recycling. The company also acknowledges the reputational problem: « The optics problem is real. AI company destroys two million books is not a headline that generates sympathy. » The suggested workaround is pure spin: reframe the destruction as digital preservation.
Buyer identities are never disclosed. The secrecy is mutual—booksellers receiving bulk orders cannot verify whether their inventory ends up in AI training pipelines. One bookseller interviewed by 404 Media described going from selling no more than 20 books per week to hundreds almost overnight, with customers ordering random selections of rare and out-of-print titles. The purchasing profile—random across subjects, heavy on out-of-print academic and literary titles, minimal interest in illustrated editions—suggests data extraction rather than reading.
Why old books command a premium
The AI industry did not turn to physical books purely out of legal caution. The choice reflects an internal crisis: the pollution of its own training data.
Models trained on synthetic text—output generated by other models—suffer a documented decline known as model collapse. Each generation trained on machine content tends to become marginally worse than the previous one. The open web is now heavily contaminated by machine-written text, making it difficult to source genuinely human-produced material at scale. ISBNdb’s blog frames it bluntly: « Physical books published before this date are structurally clean of modern poisoning tools »—a reference to techniques like Nightshade, which let authors embed invisible characters that disrupt AI training.
Pre-2022 print books sidestep the arms race entirely. They represent a fixed, human-authored corpus that cannot be quietly rewritten, contaminated by neither synthetic text nor poisoning tools. For AI labs trying to maintain data quality at scale, they represent the last reliable source of clean, curated human knowledge. The irony is that the route to preserving that knowledge digitally requires its physical destruction.
Format shifting: the legal shield and its limits
The legal framework rests on two doctrines: the first-sale doctrine and fair use. The first allows a purchaser of a physical object to resell, lend, or dispose of that object without the copyright holder’s permission. The second, under Section 107 of the Copyright Act, protects transformative uses. Anthropic argued that its format shifting—converting lawfully purchased print books into digital copies for internal, nondistributive use—fell within both. The court agreed, but only partially.
The transformation that mattered was not the conversion from print to digital per se, but what the digital copy enabled: training an AI to generate new text. The court expressly distinguished this from the pirated copies. Anthropic had also acquired over five million books from LibGen and two million more from Pirate Library Mirror, building a « central library » intended to retain « all the books in the world » indefinitely. Judge Alsup denied summary judgment on that portion, finding too many disputed factual issues. Anthropic ultimately settled the piracy claims for $1.5 billion. The message is clear: lawfully purchased books destroyed after scanning = fair use; pirated copies retained = major legal exposure.
Authors, publishers, libraries: who pays the price?
The ruling has sent ripples through multiple industries. Publishers and authors face a legal landscape that technically protects their rights in the digital domain while providing a legitimate template for lawful physical destruction. The Authors Guild, which supported the plaintiffs, maintains that AI companies should license books rather than acquire and destroy them. Multiple federal legislative proposals have stalled in committee.
For libraries and archives, the implications are more complex. University special collections become potentially strategic assets—or targets. The distinction between a book that is merely out of print and one that is genuinely rare or unique has not been tested in this context. No court has yet addressed whether the fair use analysis changes when the physical copy in question is the last known surviving edition of a work, or carries annotations, marginalia, or provenance that give it value beyond its textual content.
The crypto precedent: burn to preserve
The logic of destroying a physical object to preserve a legal position has a precedent in an unexpected place: cryptocurrency. In 2021, the Injective protocol purchased a print by the anonymous artist Banksy, burned it on camera, and sold a digital token representing the destroyed artwork—all while the blockchain recorded ownership and provenance without preserving the physical object. The art world debated whether it was a commentary on value, a publicity stunt, or simple destruction of cultural property.
Destructive book scanning has a different purpose but a similar structure: information is preserved while the physical artifact is discarded. In both cases, the digital representation is legally cleaner than the physical original—but something is lost that the digital record cannot capture. A first edition’s binding, a signed inscription, the wear pattern of a book read a hundred times: these do not survive in a PDF scan, and they are not part of the copyright analysis.
Conclusion: what the machine consumes
The practical trajectory is not difficult to extrapolate. AI companies will continue to need training data. The web will continue to be polluted by synthetic text. Pre-2022 print books will continue to represent the cleanest available source of human knowledge. And as long as the one-copy logic holds in court, the incentive to destroy those books after scanning will remain.
Congress has repeatedly considered, without ever passing, legislation that would create a licensing mechanism for AI training on copyrighted works. Without such a federal framework, the current situation—rewarding companies that purchase and destroy physical books while exposing those that download pirated copies—will continue to govern the market. The question is not whether this pattern will continue, but whether it will be regulated, taxed, or licensed before it reshapes what survives of the physical publishing record. For now, the incentive structure is clear: buy, scan, destroy, document. The rest is paper fed into machines.
Sources
- CryptoSlate — AI firms are shredding physical books because copyright law is quietly rewarding them (July 2026)
- Debevoise & Plimpton — Anthropic and Meta Decisions on Fair Use (June 2025)
- Authors Alliance — Anthropic Wins on Fair Use for Training its LLMs (June 2025)
- Ropes & Gray — From Books to Bots: Key Takeaways from the Anthropic Fair Use Decision (June 2025)
- TNW — AI firms are buying old books for slop-free training data (July 2026)
- Futurism — AI Companies Are Buying Antique Books, Ingesting Their Contents to Train Models, and Then Destroying Them (July 2026)
This article is published for informational and educational purposes only. It does not constitute investment advice. Do your own research (DYOR) before making any decisions.

