A warehouse worker cuts the spine off a hardcover book, releasing the pages, which he then runs through a top-speed scanner. The remaining book is thrown into a trash bin for shredding. A book may be a year-long work and decades in reader-finding, but destroying it takes about four seconds.
AI Generated Illustration
Over the past year, there have been reports of AI companies buying large quantities of used and rare books, scanning them on an industrial scale, and after digitization, disposing of the physical copies. The one detail that keeps appearing in the news isn't so much that these companies use books for training their AI models. After all, that was everyone's assumption. It is In reality so many of the originals are destroyed during the process that causes surprise.
The discussion, which has started, goes far beyond a few damaged paperbacks. It also concerns In reality a vast majority of human knowledge is being converted silently into training data, and whether anyone outside the companies benefits from it. For some, it's a smart, even logical, move to address the data shortage problem. For others, it's a silent kind of loss that won't be noticed until it is too late to undo it. This poses one more question below the surface: if the internet is supposed to have everything, why do AI developers need boxes of used books after all?
Why AI Developers Still Need Printed Books
The open internet is tremendous in scope but also patchy, slapdash, and fairly shallow all over the place that counts. Search engines, comment systems, and advertising blurb make up a large percentage of the thing searched. Books don't. Looking at a published book means you're at least in an author editor line editor and fact checker. That matters more than it sounds.
Language models extract regularities from all the text they see, so training on hundreds of chapters of polished prose results in a different model from a couple hundred million words of tweetlets and scanned shopping lists. Books also include topics hardly represented online, like local history, out-of-print scientific abstracts, or fifty years of the world's literature.
Much of this stuff was never converted to digitallibraries, used bookshops, millions of titles for which the whole effort never paid off.
Yet now, for a company trying to develop a system that can write and reason like a college graduate, that vast back catalogue looks less like clutter and more like a rich, unexplored resource gathering dust on a shelf. Getting that resource to a usable form Still gets fiddly.
Inside the Hidden Pipeline That Turns Books Into AI Data
The procedure that has been documented in several accounts is really not complex which makes it frightening. The books are ordered in bulk, even off flea markets, and shipped in like packs or pallets. Workers or automated machines then remove the covers of each to leave the pages lay flat. Then an incredibly fast scanner speaks through pages at hundreds per minutes where images are then processed by OCR to convert pixels in to language.
Then this raw text is through cleaning scripts that remove scanning errors, page numbers headers so on and so forth until it ends up in a training set with millions of other texts. All of that work Still does not need to be done for the book to remain intact. Quite the contrary, an undisturbed spine makes the process somewhat slower.
The flatbed, book-cradle-mounted scanner is slower than the sheet-fed scanner for that precise reason, and the reports from the companies cited indicate that speed, in the volume at which they currently operate, is tantamount to cost efficiency. The only way to get the pages return ready is to cut the spine out into single leaves first. What makes that tradeoff so painful is that the artifact is often the thing that gets lostat the same time the fully preserved text is created, the object itself is destroyed.
Why Copyright and Ownership Questions Are Becoming Harder to Ignore
The procedure that has been documented in several accounts is really not complex which makes it frightening. The books are ordered in bulk, even off flea markets, and shipped in like packs or pallets. Workers or automated machines then remove the covers of each to leave the pages lay flat. Then an incredibly fast scanner speaks through pages at hundreds per minutes where images are then processed by OCR to convert pixels in to language.
Then this raw text is through cleaning scripts that remove scanning errors, page numbers headers so on and so forth until it ends up in a training set with millions of other texts. All of that work Still does not need to be done for the book to remain intact. Quite the contrary, an undisturbed spine makes the process somewhat slower.
The flatbed, book-cradle-mounted scanner is slower than the sheet-fed scanner for that precise reason, and the reports from the companies cited indicate that speed, in the volume at which they currently operate, is tantamount to cost efficiency. The only way to get the pages return ready is to cut the spine out into single leaves first. What makes that tradeoff so painful is that the artifact is often the thing that gets lostat the same time the fully preserved text is created, the object itself is destroyed. A first edition, a note in the flyleaf, a spine scarred from a hundred commutesnone of those is going to make it through a super-high-volume, preservation-minimized scanner. The text But will. But whether that trade is even legal is a whole other, more complicated issue.
Are Libraries Losing Something That Digital Copies Cannot Replace?
Historians and archivists have argued for years that a book is more than the sum of its words. A first printing has typography choices, paper stock, binding approach that can tell you something about the time in which it was printed. Marginalia, a previous owner x2019s handwriting in the margin, a pressed flower between pages, such insight does not survive OCR.
A scanned file preserves the words. It does not preserve the object. That gap sounds pretty sweet until you think with scale: If more than a few million secondhand books are being Scan and Destroyed instead of Scan and Explore, whole segments of hardcopy culture might be disappearing at a rate no one can begin to monitor.
Edges, small fonts, infrequently published titles already are most vulnerable, because when even a small handful of copies are taken out of circulation the physical tome might as well not be in the world. This dichotomy has always characterized the process of digitisation. The innovation comes from there.
When a library digitises an old paper document, the point is to preserve it. When a company scans a book for machine learning models, the goal is to derive value from the book, and the physical print is, at best, orthogonal to that. It is that difference of purpose that this whole debate is about.
The Hidden Cost of Building More Powerful AI Models
Getting a frontier language model off the ground takes huge amounts of text, as well as the computing and electrical power needed to get it through. Firms never disclose the composition of the training corpus, and so it is almost impossible for outsiders to verify how much scanned book content has gone in, and where it has been drawn from. That opaqueness is increasingly a competitive problem in and of itself. As the easy, freely available text online gets mined and repurposed in just about every popular model, having a supply of current, quality material, via licensing arrangements, archives, or buy one get the rest free backlists, begins to become something of a strategic advantage.
The actor with the best data will matter as much as the actor with the greatest computational power. We have seen that transition in any case in licensing deals that a bunch of A.I. Companies have made with news outlets and publishers in the last couple of years. Giving up money in advance for get that content for free to the A.I. Instead of using slippery fair use rhetoric.
It is slower and more costly way to create a dataset. It is also that version of the story that does not involve demolishing anything to get the text.
What This Debate Means for the Future of Knowledge and Artificial Intelligence
Set aside the scanners and the lawsuits, and all that's left is a difficult question of who is actually in charge of controlling the reuse of human knowledge once it has been digitized and is So made it very convenient to be digitized. This question is not new and was not created by AI. Yet, AI has highlighted it so badly that libraries and archivists couldn't even come close to it being their business alone.
We do know the ways forward even though they may not exactly be simple. More comprehensive licensing agreements, greater transparency about the content of the data that is being used for training and prioritization in digitisation towards preservation, all these measures can help lower the level of friction A lot without slowing AI growth to the point of a standstill. And by the way, none of this has any bearing whatsoever on destroying a single book.
In fact, the main source of controversy has never really been a debate of whether or not AI must learn from whatever the humans have already written. It was always rather the question of whether companies that are manufacturing these AI systems have any ethical, legal or even economic responsibilities like attribution, payment or even simply attention towards the books that their AI is munching through. And so far, the answer to this question is still a mystery that will be gradually unveiled through the filing of court case and the making of licensing deal, one at a time.
