Someone Made a Dataset of One Million Bluesky Posts for 'Machine Learning Research'

https://www.404media.co/someone-made-a-dataset-of-one-million-bluesky-posts-for-machine-learning-research/

Bluesky may have said it won't use user data to train generative AI, but someone else just published a dataset of million Bluesky posts for "machine learning research". Already very popular dataset, your data may be scraped

Without paywall

44 points · 5 comments · view on lemmy.world

5 Comments

KurtVonnegut@mander.xyz · 11 pts · 1y (4 replies)

The same can and will happen with the Fediverse right?

GeneralEmergency@lemmy.world · 10 pts · 1y

Probably already happened

Viking_Hippie@lemmy.world · 1 pts · 1y (2 replies)
[ removed ]
KurtVonnegut@mander.xyz · 7 pts · 1y

I see. Probably mastodon.social gets scraped, then 🫣

ladicius@lemmy.world · 2 pts · 1y

Is that a problem for a proper scraper? Give the machine a list of domains and some hints about the relevant protocols, and then the computer runs until the end of the list.

hexagonwin@lemmy.sdf.org · 4 pts · 1y

tbh this can happen with everything now so..

i'm not sure what would be the solution, sadly.