cross-posted from: https://lemmus.org/post/24011749
AI companies are buying up and destroying tons of old books for training data
https://www.404media.co/ai-companies-are-buying-tons-of-old-books-because-theyre-free-of-ai-slop/
https://www.404media.co/ai-companies-are-buying-tons-of-old-books-because-theyre-free-of-ai-slop/
cross-posted from: https://lemmus.org/post/24011749
4 Comments
PonyOfWar@pawb.social · 9 pts · 14d
Has already been posted twice to this community within the last 24 hours.
EndlessApollo@lemmy.world · 4 pts · 14d
I remember when Google tried to do some sort of online library thing where they borrowed tons of books from libraries and uploaded them on the internet, it's good to know what that was for this whole time :D
PonyOfWar@pawb.social · 11 pts · 14d
Google books was created 22 years ago. Doubt that they had that much foresight. It was honestly a pretty cool project from Google's pre-evil days.
DonAntonioMagino@feddit.nl · 5 pts · 14d
It still is. Really useful for historical research as it contains books from the entire era of printing (so starting in the 15th century), in lots of languages, and has (though rudimentary) OCR, so you can search through texts. I used it for my thesis in university to find mentions of gothic script, as I wanted to study attitudes to this script in Dutch language works through time.
However, my thesis supervisor also mentioned to me that Google has a bad reputation in the heritage library/archival world because the scanning is done by low wage workers that have to work as quick as possible. If I remember correctly, a major Dutch library stopped lending their works to Google because of this, and they’ve made their own portal for scanned works.
Not quite so drastic as what I’ve read about these AI companies, though; pretty much just using books as raw materials. Simply criminal...