GUI browsers can ignore robots.txt, but CLI browsers cannot? Can someone explain this?

Someone brought the HTtrack app to my attention in this thread. Superficially it’s a great concept. But it did not work for me.

The author of HTtrack says they respect the robots.txt files. I’m not sure if this is my problem. But what an absurd line to draw. Why should a realtime GUI user have more privilege to view a website than a CLI user?

If anything, it should be the other way around.

  • People who lack the priviledge of having Internet at home face this nasty discrimination of being treated like a bot after they make the extra effort of commuting to a public library just to fetch a website for offline viewing later. This 2nd-classing of a demographic who is already marginalized is quite despicable.
  • People with Internet at home can schedule HTtrack to run at an off-peak time and on a narrow bandwidth put less burden on the server than the realtime GUI users who obviously hit the site mostly during peak times.
  • (update) Some people are on measured rate Internet connections that give tiny daytime quotas and generous late night quotas. Tools like HTtrack are needed to manage this. Which ultimately benefits the more privileged Internet users who have no constraints.

Update

There is a poorly worded -s0 option to ignore robots.txt. Fooled some people into thinking the tool uses robots.txt to direct the fetches.

-1 points · 0 comments · view on lemmy.world

0 Comments

No comments yet.