cross-posted from: https://lemmy.sdf.org/post/58474758
Someone brought the HTtrack app to my attention in this thread. Superficially it’s a great concept. But it did not work for me.
The author of HTtrack says they respect the
robots.txtfiles. I’m not sure if this is my problem. But what an absurd line to draw. Why should a realtime GUI user have more privilege to view a website than a CLI user?If anything, it should be the other way around.
- People who lack the priviledge of having Internet at home face this nasty discrimination of being treated like a bot after they make the extra effort of commuting to a public library just to fetch a website for offline viewing later. This 2nd-classing of a demographic who is already marginalized is quite despicable.
- People with Internet at home can schedule HTtrack to run at an off-peak time with a narrow bandwidth put less burden on the server than the realtime GUI users who obviously hit the site mostly during peak times.
- (update) Some people are on measured rate Internet connections that give tiny daytime quotas and generous late night quotas. Tools like HTtrack are needed to manage this. Which ultimately benefits the more privileged Internet users who have no constraints.
Update
There is a poorly worded -s0 option to ignore robots.txt. Fooled some people into thinking the tool uses robots.txt to direct the fetches.


I enter the library with a laptop, and list of tasks and URLs. Tasks, meaning I have to e.g. search for a PDF manual for a 2nd-hand appliance I either bought for pulled from a dumpster. Or research something. For URLs that I just need to save for later reading, I open them in FF (many tabs) and use the SingleFile extension to save them one by one. Of course that robs me of human time that I need for tasks that must be interactive. My time would be more wisely managed to have HTtrack fetching what I need in the background while I do interactive things. I sometimes stay until I get kicked out because the library is closing, in which case my needs were not all satisfied.
For Lemmy, I save posts in advance as text files and copy-paste the text into a Lemmy web client. This is also not a good use of my time but the only offline lemmy client is broken. But if there were a non-broken lemmy client for offline access, it would probably face the same discrimination by this reckless and obnoxious anti-bot movement.