"That is convenience dressed up as necessity." Destroying rare books for AI training is a choice, not a necessity, and it prioritizes speed over preservation. Non-destructive scanning offers a viable alternative, yet firms often opt for the faster, destructive method. This practice risks permanent loss for temporary gain, challenging ethical and cultural norms.
Verdict: Rarity should be preserved, not pulped.
Rarity should not be sacrificed for AI training speed
Destroying rare or scarce books to speed AI work is not necessary; it is a choice, and a bad one. Model training is the process of feeding large amounts of text into a system so it can learn patterns and generate responses, and firms are choosing the fastest intake method as if speed settled the question.
According to jpost.com, tech giants have been working on a secret project to buy millions of books and destroy them to feed new models. Booksellers also suspect AI firms are buying and then destroying rare books, which makes this a live dispute about practice, not a hypothetical fear.
That is convenience dressed up as necessity.
The actual choice
Non-destructive scanning means capturing a book’s pages without taking the book apart or ruining it, and Google patented such a method in 2009. The Internet Archive has long worked on the premise that scanning old texts takes time and attention to limit handling, which is exactly what preservation requires.
| Method | Effect on the book | Speed logic |
| Shredding for ingestion | Book is destroyed | Fastest throughput |
| Non-destructive scanning | Book remains intact | Slower, but preserves the object |
Once that option exists, the argument changes. The question is no longer whether AI needs text; it is whether firms get to erase scarce physical sources to save time, and the answer should be no.
The record already shows destruction is a chosen shortcut
Destroying rare or scarce books for model training is an avoidable choice that confuses convenience with necessity. According to arstechnica.com, a lawsuit in summer 2025 revealed that Anthropic destroyed millions of print books to train its AI models, which means the practice is not speculative or fringe but already operational at industrial scale.
That matters because the supply chain is already visible. Bulk sourcing means arranging large-volume acquisition through a supplier or broker rather than finding copies one by one, and arstechnica.com reported that ISBNdb advertised help for AI firms seeking books that way.
- Anthropic destroyed millions of print books for training, exposed by a lawsuit in summer 2025.
- ISBNdb advertised assistance to AI firms looking to source books in bulk.
- The Jerusalem Post reported that tech giants have been working on a secret project to buy millions of books and destroy them for models.
- The same reporting described the mechanism plainly: buy books, shred them, and use the text to train smarter chatbots.
- Google is involved in that buy-and-destroy project.
This is a chosen shortcut.
**Why it matters:** If you write, collect, sell, lend, or preserve books, this is not an abstract policy fight. The machinery to vacuum up physical copies and turn them into training input already exists.
Once firms can source at volume, destruction stops being an isolated bad act and becomes a procurement method. The live argument is not whether companies want more text fast; it is whether speed justifies feeding models by erasing physical copies that markets and institutions may not replace.
The best case for shredding books deserves a fair hearing
Destroying rare or scarce books for model training is an avoidable choice that confuses convenience with necessity, but the opposing case is not silly. According to arstechnica.com, a lawsuit in summer 2025 revealed that Anthropic destroyed millions of print books to train its AI models, which shows why firms think in terms of industrial throughput rather than one-by-one care. A corpus is the large body of text used to train a model.

That pressure is real.
The strongest case on the other side
The steelman starts with speed and substitution:
- many print copies are not unique, so buying and cutting them apart looks like using surplus rather than erasing culture
- a training corpus has to be vast and assembled fast, because delay means weaker products and lost market position
- preservation-first scanning is slower by design, because careful handling takes time and attention, as institutions that scan old texts already know
Claim: If firms need huge text collections quickly, destructive scanning is the only practical way to compete at market speed.
Counterargument: Google patented a non-destructive book-scanning technology in 2009, so there are alternatives to cutting books apart.
Rebuttal: An alternative existing on paper does not erase the business incentive to choose the faster workflow, but that only proves destruction is a shortcut, not a necessity. Once the shortcut reaches scarce material, the speed argument fails because the loss is permanent while the gain is only convenience.
Speed is not a licence to erase what cannot be replaced
Rarity should not be sacrificed for AI training speed, because once a scarce work is destroyed the gain is temporary and the loss is permanent. Ars Technica reported that booksellers suspect AI firms are buying and then destroying rare books, which matters because the practice treats replacement as if it were guaranteed when rarity means it is not.
A cultural institution is an organization that keeps and passes on shared knowledge, records, and values, such as libraries, archives, museums, and universities. Those institutions and our ethics were shaped on a narrow class of examples, according to ucr.edu, so they are already selective in what they notice and keep.
That is exactly why speed is not a licence to erase what cannot be replaced.
When preservation is already narrow, governance has to be stricter, not looser.
- We keep only part of the record.
- We judge value through inherited norms.
- We are using stronger tools while wisdom lags.
That combination changes the practical calculus. If a faster pipeline destroys rare material for a modest training gain, it converts uncertainty into certainty: the book is gone, the benefit is contestable, and the option to preserve first has been discarded. Good governance does not accept irreversible loss for marginal gain when the loss falls on the public record and the gain accrues to a private build schedule. The same basic rule applies wherever power outruns judgment: a single mistake by a sufficiently powerful person or group can make conditions far worse for everyone, so institutions should avoid choices that narrow our margin for error.
Rarity should be preserved, not pulped
No: rare books should not be destroyed to feed AI, because once non-destructive scanning exists, shredding scarcity is not necessity but convenience purchased with irreversible public loss. The answer flips only if a copy is genuinely non-rare, fully replaceable, and destruction is the only way to obtain text that cannot be captured without harm; absent that narrow condition, speed is not a licence to erase cultural record for a private training schedule, and rarity should be preserved, not pulped.
Sources
- 1. [Why AI companies are buying millions of books to destroy them | The Jerusalem Post]( https://www.jpost.com/business-and-innovation/tech-and-start-ups/article-904336)
