TechTonic Times
Feel the Pulse of Progress
Artificial Intelligence & Data

Millions of Books Were Destroyed to Train AI. Now Questions Grow Over the Cost of Building Large Language Models

A warehouse worker cuts the spine off a hardcover book, releasing the pages, which he then runs through a top-speed scanner. The remaining book is thrown into a trash bin for shredding. A book may be a year-long work and decades in reader-finding, but destroying it takes about four seconds.

An industrial scanner digitizes a hardcover book moments before it enters a paper shredder, illustrating reports that books are being scanned for AI training and then destroyed.

AI Generated Illustration

Over the past year, there have been reports of AI companies buying large quantities of used and rare books, scanning them on an industrial scale, and after digitization, disposing of the physical copies. The one detail that keeps appearing in the news isn't so much that these companies use books for training their AI models. After all, that was everyone's assumption. It is In reality so many of the originals are destroyed during the process that causes surprise.

The discussion, which has started, goes far beyond a few damaged paperbacks. It also concerns In reality a vast majority of human knowledge is being converted silently into training data, and whether anyone outside the companies benefits from it. For some, it's a smart, even logical, move to address the data shortage problem. For others, it's a silent kind of loss that won't be noticed until it is too late to undo it. This poses one more question below the surface: if the internet is supposed to have everything, why do AI developers need boxes of used books after all?

Why AI Developers Still Need Printed Books

The open internet is tremendous in scope but also patchy, slapdash, and fairly shallow all over the place that counts. Search engines, comment systems, and advertising blurb make up a large percentage of the thing searched. Books don't. Looking at a published book means you're at least in an author editor line editor and fact checker. That matters more than it sounds.

Language models extract regularities from all the text they see, so training on hundreds of chapters of polished prose results in a different model from a couple hundred million words of tweetlets and scanned shopping lists. Books also include topics hardly represented online, like local history, out-of-print scientific abstracts, or fifty years of the world's literature.

Much of this stuff was never converted to digitallibraries, used bookshops, millions of titles for which the whole effort never paid off.

Yet now, for a company trying to develop a system that can write and reason like a college graduate, that vast back catalogue looks less like clutter and more like a rich, unexplored resource gathering dust on a shelf. Getting that resource to a usable form Still gets fiddly.

Inside the Hidden Pipeline That Turns Books Into AI Data

The procedure that has been documented in several accounts is really not complex which makes it frightening. The books are ordered in bulk, even off flea markets, and shipped in like packs or pallets. Workers or automated machines then remove the covers of each to leave the pages lay flat. Then an incredibly fast scanner speaks through pages at hundreds per minutes where images are then processed by OCR to convert pixels in to language.

Then this raw text is through cleaning scripts that remove scanning errors, page numbers headers so on and so forth until it ends up in a training set with millions of other texts. All of that work Still does not need to be done for the book to remain intact. Quite the contrary, an undisturbed spine makes the process somewhat slower.

The flatbed, book-cradle-mounted scanner is slower than the sheet-fed scanner for that precise reason, and the reports from the companies cited indicate that speed, in the volume at which they currently operate, is tantamount to cost efficiency. The only way to get the pages return ready is to cut the spine out into single leaves first. What makes that tradeoff so painful is that the artifact is often the thing that gets lostat the same time the fully preserved text is created, the object itself is destroyed.

Why Copyright and Ownership Questions Are Becoming Harder to Ignore

The procedure that has been documented in several accounts is really not complex which makes it frightening. The books are ordered in bulk, even off flea markets, and shipped in like packs or pallets. Workers or automated machines then remove the covers of each to leave the pages lay flat. Then an incredibly fast scanner speaks through pages at hundreds per minutes where images are then processed by OCR to convert pixels in to language.

Then this raw text is through cleaning scripts that remove scanning errors, page numbers headers so on and so forth until it ends up in a training set with millions of other texts. All of that work Still does not need to be done for the book to remain intact. Quite the contrary, an undisturbed spine makes the process somewhat slower.

The flatbed, book-cradle-mounted scanner is slower than the sheet-fed scanner for that precise reason, and the reports from the companies cited indicate that speed, in the volume at which they currently operate, is tantamount to cost efficiency. The only way to get the pages return ready is to cut the spine out into single leaves first. What makes that tradeoff so painful is that the artifact is often the thing that gets lostat the same time the fully preserved text is created, the object itself is destroyed. A first edition, a note in the flyleaf, a spine scarred from a hundred commutesnone of those is going to make it through a super-high-volume, preservation-minimized scanner. The text But will. But whether that trade is even legal is a whole other, more complicated issue.

Are Libraries Losing Something That Digital Copies Cannot Replace?

Historians and archivists have argued for years that a book is more than the sum of its words. A first printing has typography choices, paper stock, binding approach that can tell you something about the time in which it was printed. Marginalia, a previous owner x2019s handwriting in the margin, a pressed flower between pages, such insight does not survive OCR.

A scanned file preserves the words. It does not preserve the object. That gap sounds pretty sweet until you think with scale: If more than a few million secondhand books are being Scan and Destroyed instead of Scan and Explore, whole segments of hardcopy culture might be disappearing at a rate no one can begin to monitor.

Edges, small fonts, infrequently published titles already are most vulnerable, because when even a small handful of copies are taken out of circulation the physical tome might as well not be in the world. This dichotomy has always characterized the process of digitisation. The innovation comes from there.

When a library digitises an old paper document, the point is to preserve it. When a company scans a book for machine learning models, the goal is to derive value from the book, and the physical print is, at best, orthogonal to that. It is that difference of purpose that this whole debate is about.

The Hidden Cost of Building More Powerful AI Models

Getting a frontier language model off the ground takes huge amounts of text, as well as the computing and electrical power needed to get it through. Firms never disclose the composition of the training corpus, and so it is almost impossible for outsiders to verify how much scanned book content has gone in, and where it has been drawn from. That opaqueness is increasingly a competitive problem in and of itself. As the easy, freely available text online gets mined and repurposed in just about every popular model, having a supply of current, quality material, via licensing arrangements, archives, or buy one get the rest free backlists, begins to become something of a strategic advantage.

The actor with the best data will matter as much as the actor with the greatest computational power. We have seen that transition in any case in licensing deals that a bunch of A.I. Companies have made with news outlets and publishers in the last couple of years. Giving up money in advance for get that content for free to the A.I. Instead of using slippery fair use rhetoric.

It is slower and more costly way to create a dataset. It is also that version of the story that does not involve demolishing anything to get the text.

What This Debate Means for the Future of Knowledge and Artificial Intelligence

Set aside the scanners and the lawsuits, and all that's left is a difficult question of who is actually in charge of controlling the reuse of human knowledge once it has been digitized and is So made it very convenient to be digitized. This question is not new and was not created by AI. Yet, AI has highlighted it so badly that libraries and archivists couldn't even come close to it being their business alone.

We do know the ways forward even though they may not exactly be simple. More comprehensive licensing agreements, greater transparency about the content of the data that is being used for training and prioritization in digitisation towards preservation, all these measures can help lower the level of friction A lot without slowing AI growth to the point of a standstill. And by the way, none of this has any bearing whatsoever on destroying a single book.

In fact, the main source of controversy has never really been a debate of whether or not AI must learn from whatever the humans have already written. It was always rather the question of whether companies that are manufacturing these AI systems have any ethical, legal or even economic responsibilities like attribution, payment or even simply attention towards the books that their AI is munching through. And so far, the answer to this question is still a mystery that will be gradually unveiled through the filing of court case and the making of licensing deal, one at a time.

Important Note

This article is based on information from publicly available sources, including official announcements, research publications, and reputable news outlets available at the time of writing. While every effort has been made to verify the accuracy of the information, errors or omissions may still occur. The content is provided for informational purposes only and should not be considered professional medical, legal, financial, or technical advice. Readers are encouraged to consult original sources and qualified professionals before making decisions based on the information presented.

Spread the Word

About the Author

Mir Mushfikur Rahman

Mir Mushfikur Rahman

Founder & Editor

Covering Breakthrough Technologies, Medical Innovations, Daily Science And The Future Of Science. Dedicated To Making Complex Tech Accessible To Everyone.

Editor's Picks

Frequently Asked Questions

AI developers destroy physical books to achieve the high-speed scanning required for mass data extraction. Cutting the spines allows sheet-fed scanners to process hundreds of pages per minute, providing the polished, fact-checked text that large language models need, though it sacrifices the physical artifact.
The legality remains highly contested. While companies often claim fair use for training data, authors and publishers argue it infringes on copyright. The lack of transparency in training datasets has sparked ongoing lawsuits and pushed the industry toward formal licensing agreements to avoid legal risks.
The open internet is often filled with shallow, unverified, and noisy content like social media posts and ads. Books provide high-quality, deeply researched, and fact-checked prose, along with niche topics like local history and out-of-print science, which are essential for training advanced reasoning capabilities.
When books are shredded after scanning, historians lose valuable physical artifacts. Elements like marginalia, first-edition typography, paper stock, and pressed flowers provide historical context that OCR text extraction cannot capture, leading to a silent loss of hardcopy culture and archival heritage.
Publishers and authors are increasingly demanding transparency and compensation for the use of their work in AI training. Instead of relying on disputed fair use claims, many are pushing for comprehensive licensing agreements and ethical data practices to ensure creators benefit from their intellectual property.