archivist

u/archivist@lemm.ee
34 posts · 10 comments

Recent posts

Recent comments

In case you haven't looked into it yourself yet...

ArchiveTeam are independent from IA, but their stuff mostly does end up uploaded into the Wayback Machine. Storage space (like yours) isn't usually what they are looking for, but rather the internet bandwidth and "virgin" IP address of aforementioned "warriors" running their code to scrape different websites, and then uploading the results to AT's servers, where they are collected and eventually uploaded again to IA.

Check out https://tracker.archiveteam.org/ for current projects

ArchiveTeam are "planning" to save everything(?), but they should have started like two months ago, if they have any hope of doing that. Dunno if it's being worked on at all. Plus committing to a multiple petabyte project I assume takes some doing.

I've scarcely visited the site myself, but I looked around for stuff of interest to me and bagged them. yt-dlp works just fine!

After a while, blog posts started returning a 403 error, then later 410. Images and javascript files remained downloadable for longer, but the JS files started returning 410 after a while as well. Now, only images are available, and the known ones are slowly being archived as long as they are downloadable.