I could really use a commandline mechanism for saving a webpage locally for offline viewing later. Saving the page in a single file is ideal.

“singlefile”, the browser extension

For GUI users there is a quite useful extension (single-file). That can be run from the CLI by installing single-file-cli. But I am not enthusiastic about the bloaty option. You must first install some obscure package manager called “Deno”, which does not exist in the official Debian repositories (thus has not proven to be worthy for mainstream public use, and also complicates every Debian upgrade). Then it uses JavaScript to instrument a bloaty GUI browser. Not great.

Worth noting that there is Org-mode integration but this does not get around the the non-Debian dependency on Deno.

wget

A page can be fetched using this command:

$ wget -U 'stopDiscriminatingAgainstWg3t' -c -P "$dir" -E -H -k -K --xattr -p "$url"
$ printf '%s' "$url" > "$dir"/url.txt

The printf is needed because wget just dumps a big tree of files with no indication of which file must be opened to render the whole page. Indeed it sucks. Saving the URL gives us a fighting chance at locating the file but it’s still a pain. And from there, the files are still scattered on your hard drive.

There is a non-Debian tool called HTMLArk. Again, I am not thrilled about non-Debian stuff. But it claims to be able to produce a single HTML file from a scattered collection of files. I have not tried it but in principle it can be used to tidy up the mess dumped by wget.

ArchiveBox and Grunt-inline

ArchiveBox looks interesting. But it’s non-Debian so not exactly spot-on. Same issue with Grunt-inline. There is some grunt stuff in the Debian repos but not this tool specifically.

webpages2html

This is a python script but the script and its dependencies are non-debian. Requires using pip which is a disaster of a tool that is barely suitable for developers but hardly suitable for the end users who end up getting pushed into it b/c there is no proper pkg manager.

What else?

Any decent option I have overlooked?

  • evenwichtOPM
    link
    fedilink
    arrow-up
    1
    ·
    edit-2
    4 days ago

    I usually want to archive a page that can be faithfully reproduced, sometimes for borderline forensic purposes. I usually don’t have time during a browsing session to inspect the results of a saved page for quality. If I pop into a public hotspot I want to save and go.

    But later on when I’m reading and managing the offline content, I might decide the graphics on a page are unwanted. So perhaps it would be useful to do w3m -dump "$my_local_GUI_page.html" > lean_capture.txt && rm $my_local_GUI_page.html". I’ll have to try that.

    (update) I works to run w3m -dump "$my_local_GUI_page.html" > lean_capture.txt but strangely the file:// scheme is not understood. E.g. this does not work: w3m -dump file://./"$my_local_GUI_page.html".

    (update 2) If the webpage is MIME-formatted, w3m does not handle it and just dumps the base64 blobs.