AI companies opt for the destructive route... The result is the destruction of millions of books.

https://www.tomshardware.com/tech-industry/artificial-intelligence/ai-companies-are-reportedly-shredding-millions-of-books-to-train-models-tech-giants-outsource-to-middlemen-to-secretly-buy-up-books-for-training-material

During the Anthropic lawsuit, Tom Harvey, who previously participated in the creation of Google Books before leading Anthropic's "Project Panama" digitalization project, confirmed that the AI firm hired several document scanning companies. Datamation Information Services, which offers high-volume, non-destructive, and destructive book scanning services, was one of them. The former method employs different tools, like overhead scanners, flatbed scanners, or V-shaped imaging systems. The latter method, on the other hand, would have personnel gut the books and feed the individual pages into a high-speed industrial scanner. Logically, AI companies opt for the destructive route since it is more efficient and lower-cost. The result is the destruction of millions of books.

Obviously, printed books represent a treasure trove of information for AI. However, many debate the ethics of removing books from circulation since it is uncertain whether AI companies filter the rare or even out-of-print books from the common titles during digitalization. The other major issue is that scanned books go directly into a private database to train AI, which the general public does not have access to. True, we will have smarter AI, but at the cost of the information not being available to future generations.

Source [web-archived; 2026-07-28_18-53-43Z]

Related: AI Companies Are Buying Tons of Old Books... (...documents about its plan to scan millions of books... to destroy the books in the process...) [web-archived; 2026-07-21_14-58-43_Z]

171 points · 10 comments · view on lemmy.world

10 Comments

from_D4rkness@lemmy.world · 39 pts · 16d (2 replies)

This will be seen as the new "burning of the library of alexandria" event, if it will even be remembered.

artwork@lemmy.world · 19 pts · 16d (1 reply)

Ineffably sorrowful... no words may express it...

We are burning several libraries of Alexandria to write slightly better LinkedIn posts...

Source [2026-07-29 03-50-41Z]

wheezy@lemmy.ml · 3 pts · 16d

The cope is that the post are better. AI is rotting the brains of all the people that use it for writing. Writing is a fundamental exercise for our brains; that all of us should be using in at least a small part every day. Hell, even social media comments are working the muscles of your brain to an extend. When you stop using that muscle it's going to weaken. Remember that for everything you use AI for.

And no, writing prompts are not working that muscle. You're writing a prompt specifically because it's less cognitively straining to do so. You're using 2lb dumbbells when you were previously using 50lbs.

gravitas_deficiency@sh.itjust.works · 28 pts · 16d

I guess one of the more important questions is “do they even bother to attempt to dedupe via ISBN or anything like that? Because if they’re getting rare editions alongside common editions, and then not bothering to check that they’ve already got the content and shredding both…

Let’s be honest: they don’t give a shit, and they’re probably doing exactly what I’m describing.

botbot@feddit.org · 26 pts · 16d

This should be illegal

fubarx@lemmy.world · 7 pts · 16d

If it's any consolation, AI training on printed words (or music, art, or video) is just as useful as tasting a stew so you can learn farming.

4am@lemmy.zip · 3 pts · 16d

Just like tech: if you can’t get it anywhere else, you’ll have to rent it from them forever

Hamartia@lemmy.world · -1 pts · 16d (2 replies)

From an archival point of view, nothing that they're destroying doesn't exist in numerous libraries. Yes, destroying rare books isn't good but archives the world over will hold copies of these tomes.

If this was an actual concern archivists the world over would be shouting about it. In my experience they are rabid protectors of our material heritage. In reality they see digitisation as a significant part of the preservation of that heritage.

As an example: paper production went through a crisis 1800 - 1850; it used to be made from discarded cotton cloth (and was of high quality); but by 1800 not enough old cloth could be sourced to satisfy demand; so they started to make paper from pulped wood; which was very bad as it contained lignin which is acidic; but it took years to figure this out; in western archives are lots of books and newspapers turning into confetti from that period; we can save some material from that period by washing out the acidity and lining the pages with very thin, very strong Japanese tissue; it just isn't economically viable to save all extant copies of some mass produced book/newspaper; better to preserve some select holdings in special repository libraries and digitise it for posterity.

There are way more important issues to get riled around the arseholes developing AI and the data centres they demand. This is just rage baiting.

Get_Off_My_WLAN@fedia.io · 6 pts · 16d (1 reply)

The problem is, among basically three combinations of choices they can make,

  1. keep the data for themselves, but digitize the rare books without destroying the original
  2. destroy the rare books, but digitize the data and make it publically available
  3. digitize the rare books while destroying the original and keeping the data to themselves

they're choosing the worst option possible for humanity.

Hamartia@lemmy.world · 3 pts · 15d

None of that is as near a worrying issue as the environmental destruction from their resource hungry data centers; the surveillance systems they enhance; or the wars they help prosecute.

My point in the first post is that your idea of what constitutes a rare book will be fairly different to how an actual archivist sees it. It is very very unlikely that the tech bros can get their hands on books that are not held in multiple museums/libraries/archives. That has been the collection policy since a couple of very important archives got flooded in the mid C20th.

Even the breaking up of books isn't unusual. All the archives taking part in digitising projects will have a program where old books with overly tight text blocks are broken up for the scanning process then rebound (just as the majority of very old books have had done multiple times over their lifetime) afterwards. Most modern books with adhesive bindings, and sufficient extant copies, will just get thrown in the recycling after scanning as they are not as archively important enough to warrant the outlay of resources to rebind.

In an ideal world we wouldn't be destroying any books but preserving books has an embodied cost, financially and environmently, so archivists have to deside how much is ring fenced. They have been doing this for decades now so the tech bros won't be able to get their hands on anything the archivists haven't already assessed as surplus to requirements.