# Y2Z/monolith: Save a Web Page as One HTML File

> monolith is a Rust CLI and library that inlines CSS, images, fonts and JavaScript into a single HTML5 document. It is built for archiving, not for scraping, and it cannot execute JavaScript.

**Y2Z/monolith** — ⬛️ CLI tool and library for saving complete web pages as a single HTML file

- Repository: https://github.com/Y2Z/monolith
- Website: https://crates.io/crates/monolith
- Stars: 15,504 · Forks: 474
- Language: Rust
- License: CC0-1.0
- Published: 2026-09-21 · Updated: 2026-09-21 · Language: en
- Canonical page: https://hysenlabs.com/projects/y2z-monolith

## What monolith solves, and who actually needs it

The browser's "Save page as" produces a folder of loose assets with relative links that break the moment you move or rename anything. monolith produces one file. The README frames the tool as "a data hoarder's dream come true: bundle any web page into a single HTML file," and that framing is accurate: the output is an HTML5 document with CSS, images and JavaScript embedded as data URLs, so a browser can render it with no network connection at all.

The people this fits are archivists, support engineers capturing a bug report before a site changes, and anyone who keeps a personal reading archive on disk. It also fits build pipelines that need a deterministic, self-contained artifact from a URL. It does not fit anyone who needs to extract structured data from a page, and it does not fit anyone who needs the page as a screenshot or PDF. The output is a document, not a dataset.

## How the inlining pipeline works

The Cargo.toml lists html5ever for "all things DOM" and markup5ever_rcdom for manipulating that DOM, which tells you the shape of the pipeline: monolith parses the target document into a DOM, walks it, fetches each referenced asset, base64-encodes the bytes, and rewrites the reference in place. cssparser handles stylesheets, url and percent-encoding handle link resolution, and encoding_rs handles charset conversion so the saved document is not mangled when the source is not UTF-8.

There is no browser engine in that list. The README states plainly that "Monolith doesn't feature a JavaScript engine, hence websites that retrieve and display data after initial load may require usage of additional tools." That single sentence explains most of the tool's behaviour: it archives the HTML the server sent, plus whatever assets that HTML references. Anything a client-side framework would have injected later simply is not there.

The dependency list also includes redb, described as being "for on-disk caching of remote assets," and sha2, used "for calculating checksums during integrity checks." So repeated runs against the same page can reuse cached assets, and integrity attributes are computed rather than copied blindly. Neither behaviour is documented in the README beyond those one-line comments in Cargo.toml, which is a gap: there is no published description of where the cache lives, how it is invalidated, or how to clear it.

## Installing monolith and saving your first page

The README lists a long set of package managers. On any platform with a Rust toolchain, cargo is the shortest path:

```bash
cargo install monolith
```

On macOS or GNU/Linux, Homebrew is available, and on Windows there are Chocolatey, Scoop and Winget entries. Pick whichever matches the machine you are on; the README gives the exact command for each.

Once installed, the README's own usage example writes the page to a filename built from the page title and a timestamp:

```bash
monolith https://lyrics.github.io/db/P/Portishead/Dummy/Roads/ -o %title%.%timestamp%.html
```

After that runs you should have a single .html file in the current directory whose name contains the page title. Open it in a browser with the network disconnected; if the page renders, the inlining worked.

The second documented example reads a document from standard input and writes it to standard output, which is how you feed monolith a page you already rendered elsewhere:

```bash
cat some-site-page.html | monolith -aIiFfcMv -b https://some.site/ - > some-site-page-with-assets.html
```

That flag cluster is worth unpacking because it is the actual configuration surface. It excludes audio sources (-a), removes images (-i), isolates the document (-I), omits frames (-f), excludes CSS (-c), excludes videos (-v), and sets a custom base URL (-b) so relative links in the piped HTML resolve correctly. If you pipe HTML in without -b, relative asset URLs have no origin to resolve against.

## Controlling which domains get fetched

Options -d and -B are the allowlist and blocklist, and the README's examples show them used together with -I (isolate the document) to produce a page that only pulls from named hosts:

```bash
monolith -I -d example.com -d www.example.com https://example.com -o example-only.html
```

The second example shows the opposite intent, blocking analytics and ad hosts while allowing a CDN:

```bash
monolith -I -B -d .googleusercontent.com -d googleanalytics.com -d .google.com https://example.com -o example-no-ads.html
```

Note the leading dot on .googleusercontent.com and .google.com. That is a subdomain match, and it is the kind of detail that decides whether your filter does what you think. The README does not spell out the matching rules for -d and -B beyond these examples, so if your archive depends on a precise allowlist, test it against a page with assets on several subdomains before you trust the result. A filter that silently drops an asset leaves you with a saved page that looks fine until the network is gone and a stylesheet is missing.

## Where monolith is the wrong tool

The JavaScript limitation is the one that matters. A page whose content arrives via a client-side fetch after load will be saved as an empty shell. The README does not hide this; it points at a workaround, using Chromium as a pre-processor:

```bash
chromium --headless --window-size=1920,1080 --run-all-compositor-stages-before-draw --virtual-time-budget=9000 --incognito --dump-dom https://github.com | monolith - -I -b https://github
```

That pipeline is two tools, not one, and it inherits Chromium's cost and its own flags. The --virtual-time-budget value is a guess about how long the page needs; too low and you archive a half-rendered document. The README presents this as an example, not as a supported integration, so treat the flag values as starting points to adjust rather than defaults.

The second limitation is scale. There is no documented crawl mode, no queue, no sitemap handling and no resume. monolith takes a URL or a document and writes a file. If you need to archive a thousand pages, you are writing the loop yourself, and you will be dealing with the on-disk asset cache whose behaviour is described only by a comment in Cargo.toml.

The third is fidelity. Embedding assets as data URLs makes files large. A media-heavy page produces a correspondingly large HTML file, and the README offers no size guidance. For text-heavy documentation pages this is a non-issue; for a page with a dozen high-resolution images it may be.

## monolith compared with wget and with headless browsers

The README makes one direct comparison: against wget -mpk. The difference it names is that monolith "embeds all assets as data URLs and therefore lets browsers render the saved page exactly the way it was on the Internet, even when no network connection is available." wget's mirror mode writes a directory tree with linked files. Both preserve the page, but only one survives being emailed as an attachment or dropped into a single-file archive.

Against a headless browser used as an archiver, the difference runs the other way. A headless browser executes JavaScript and can therefore capture post-load content, but its output is a rendered artifact rather than an editable HTML document with the original markup intact. monolith keeps the source document and inlines what it references. The README's own recommended pattern combines the two, which is a fair summary of the trade-off: use the browser to produce the DOM, use monolith to make it portable.

One output detail worth knowing: the -m flag writes MHTML instead of HTML. That is the format some email clients and older tooling expect. The README lists the flag but does not document how MHTML output differs in practice, so if your downstream consumer needs MHTML, verify it against that consumer before building a workflow on it.

## Maintenance, licence and upgrade cost

The repository is not archived, and the last push was on 2026-05-25. Releases are tagged and versioned: v2.10.1 on 2025-03-30, v2.10.0 on 2025-03-24 and v2.9.0 on 2025-03-11. Cargo.toml carries version 2.11.0, so the manifest is ahead of the most recent release listed. Pre-built binaries ship with each release for Windows, GNU/Linux and non-standard CPU architectures, according to the README, which means you are not forced to compile unless you want the library.

Upgrade cost is low in the normal case. The Makefile's install target runs cargo install --force --locked --path ., and the build target runs cargo build --locked, so lockfile-pinned builds are the intended path and dependency drift is not forced on you. The Cargo.toml pins every dependency with an exact version (for example base64 = "=0.22.1"), which keeps builds reproducible but means a security fix in a dependency requires a monolith release or a manual bump rather than a floating range resolving on its own.

The licence is CC0-1.0, which the repository states in both Cargo.toml and the LICENSE file. CC0 is a public-domain dedication rather than a permissive software licence in the MIT or Apache sense, and some organisations have policies about which licence identifiers they accept for dependencies. That is a policy question for your own legal reviewers, not something this article can settle; the relevant fact is simply that the identifier is CC0-1.0 and not a more common software licence.

## Conclusion

Adopt monolith if you archive pages you have already rendered or that ship static HTML, and if a single self-contained file matters more than fidelity on JavaScript-heavy sites. Do not adopt it as a crawler, a headless browser, or a screenshot service; it has no JavaScript engine, and the README says so. Before committing, check the -d and -B domain filters against your own targets, confirm the -m MHTML output path works for your viewers, and read the CC0-1.0 licence text rather than assuming it matches your organisation's policy.

## FAQ

### How do I install monolith?

The README lists many routes. The cross-platform one is cargo install monolith; Homebrew, Chocolatey, Scoop, Winget, Snapcraft, Pacman, apk and others are also documented, and each release ships pre-built binaries.

### How do I use monolith to save a page?

Pass the URL and an output path: monolith https://lyrics.github.io/db/P/Portishead/Dummy/Roads/ -o %title%.%timestamp%.html. The README also shows reading HTML from standard input and writing the bundled result to standard output.

### Does monolith run JavaScript when saving a page?

No. The README states that monolith does not feature a JavaScript engine, so pages that fetch and display data after initial load may need another tool. The README suggests using Chromium headless with --dump-dom as a pre-processor and piping the result into monolith.

### Can monolith save a page as MHTML instead of HTML?

Yes. The -m option outputs in MHTML format instead of HTML. The README lists the flag but does not describe how the two outputs differ in practice.

### How does monolith differ from wget -mpk?

The README says monolith embeds all assets as data URLs, so the saved page renders in a browser with no network connection, whereas wget's mirror mode leaves assets as separate linked files.

## Sources

- [License: CC0-1.0](https://github.com/Y2Z/monolith/blob/master/LICENSE)
- [Project website](https://crates.io/crates/monolith)
- [README](https://github.com/Y2Z/monolith/blob/master/README.md)
- [Releases](https://github.com/Y2Z/monolith/releases)
- [Y2Z/monolith on GitHub](https://github.com/Y2Z/monolith)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/y2z-monolith
