Just set a cron job up to scrape and md5 all of the documents every 20m or so; if the MD5s match, discard. If they do not, save. Then you can build a timeline of releases, redactions, edits, and fuckery.
UPDATE: As of 2:17pm ET on 1/30/26, the document is once again appearing on the DOJ website. We are looking through to see if additional redactions were made.
7 Comments
gravitas_deficiency@sh.itjust.works · 55 pts · 235d
Just set a cron job up to scrape and md5 all of the documents every 20m or so; if the MD5s match, discard. If they do not, save. Then you can build a timeline of releases, redactions, edits, and fuckery.
ltxrtquq@lemmy.ml · 21 pts · 235d
I don't know how to do any of that, but it sure sounds like you do
herseycokguzelolacak@lemmy.ml · 4 pts · 235d
Please share the whole dataset so we can clone it!
Triumph@fedia.io · 35 pts · 235d
"Oops."
KingGordon@lemmy.world · 5 pts · 235d
Vile.
kent_eh@lemmy.ca · 4 pts · 235d
But also very predictable.
HydraBenny@lemmy.world · 2 pts · 233d
There's this one too
https://www.justice.gov/epstein/files/DataSet%2010/EFTA01660651.pdf