# web-archive declares ISC in its manifest and GPL-3.0 in its repository, and the Docker path runs a local Workers runtime

> A self-hosted page archiver built from three pieces, a browser extension that saves a page as a single HTML file, a Hono server on Cloudflare D1 and R2, and a client that queries and displays it. The interesting corners are the licence conflict, a manifest version older than its tag, and a Docker image that serves the worker locally and keeps the database in a volume.

**Ray-D-Song/web-archive** — Selfhost web archiving and sharing service.

- Repository: https://github.com/Ray-D-Song/web-archive
- Website: https://web-archive-docs.pages.dev/
- Stars: 933 · Forks: 288
- Language: TypeScript
- License: GPL-3.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/ray-d-song-web-archive

## The two licence statements cannot both be right

The repository is recorded as GPL-3.0. The package manifest in the same tree says ISC. Nothing in the readme or the linked documentation resolves it, so anyone planning to fork, redistribute or embed this code has to read the licence file in the repository and decide for themselves, and anyone planning to publish a derivative is the person who has to be careful. The same manifest is out of step with the releases in a second way: its version field reads 0.1.1-alpha.1, while the newest tag is v0.2.0, published on 2026-04-27, with v0.1.3 and v0.1.2 sitting back in December 2024. So the number in the file is older than the number on the tag, which is the opposite of the usual direction and means the file is not a reliable indicator of what you deployed.

## Capture is one HTML file, uploaded by an extension to an address you type in

The system is three parts, and the readme is precise about the seam between them. The browser extension saves the current page as a single HTML file and uploads it to your server. The server receives that file and stores it in two places, a database and a storage bucket. The web client queries the file and displays it. Authentication is described as entering the service address and a key in the extension after deployment, which means the trust model is a shared string typed into a browser extension, and everything downstream, the D1 rows and the R2 objects, sits behind it. The feature list covers archiving, search and sharing, folder classification, mobile adaptation, reading mode and automatic tag classification by a model. Nothing in the description crawls a site, and no separate asset fetching is described, so what you keep is what the extension embedded in that one file.

## The Docker image runs the worker locally and keeps its data in a volume

The single command offered for local use is:

```bash
docker run -d -p 8787:8787 -v web-archive-data:/app/service/.wrangler state ghcr.io/ray-d-song/web-archive:latest
```

What the image actually is becomes clear from the Dockerfile. The runtime stage is a Node 24 Alpine image with a package called node-cf-worker installed globally, and the entry point is that binary pointed at a wrangler configuration file on port 8787. It is a local Workers runtime, so the same worker code runs locally and on Cloudflare, and the directory `.wrangler/state` is declared as a volume, which is where a local D1 and R2 keep their data. That volume is the whole archive: a named volume in the command above, so removing the container keeps your saved pages and a careless prune does not. The image is referenced by its latest tag, which for a service holding your own saved content is the tag to avoid.

## The build stage installs dependencies before it copies the plugin manifest

The build stage copies the workspace file, the lockfile and the workspace definition, then the package manifests of three packages, the server, the web client and the shared code, and only then runs a frozen lockfile install. The full source tree is copied afterwards, which is the moment the plugin package's own manifest finally arrives. The workspace is declared as packages with a wildcard, so at install time one of its declared members is not on disk. Whether a frozen install tolerates that or complains depends on how the lockfile records that member, and nothing in the readme or the linked deployment guide says. The ordering is unusual enough to be worth checking on your first build rather than assuming it works. The other line in that chain is `node scripts/build_clear.mjs`, a build step that runs after the server and web builds and is named as if it strips something out of the output.

## Four packages, one command per target, and docs built with VitePress

The repository is a pnpm workspace, which the root manifest declares alongside a pinned package manager version carrying a full integrity hash, so a different pnpm on your machine is refused rather than trusted. Everything is filtered per package: local initialisation runs against the server, development commands exist for the web client, the server and the plugin, and the service build chains the server build, the web build and the clear script. The plugin has two build targets, one general and one specifically for Firefox, which matches the fact that the extension is published in two stores. Documentation has its own three commands for developing, building and previewing a VitePress site. Linting runs through a flat ESLint configuration at the root, a TypeScript base configuration sits alongside it, and husky is wired as the prepare script, so a commit hook is installed as a side effect of installing dependencies.

## The operational documentation lives on a Pages site, not in the repository

The readme is short because most of it is a directory. It opens with two language links, a Chinese readme and the English one, and everything else is a summary plus two deployment routes. The first is Cloudflare, marked recommended, pointing at a deploy document on a separate site. The second is the Docker command above, pointing at the same document with an anchor for the local deployment section. The homepage recorded for the repository is that documentation site, hosted on Cloudflare Pages, and the Docker path is the only route described without leaving the readme. So a self-hoster ends up on a documentation site for the two things that actually matter, how to create the database and bucket, and what key to type into the extension, and the site is a separate deployment from the code it describes.

## What one file per page means for search, tags and reading mode

Three of the advertised features are downstream of a single decision: the extension produces one HTML file per saved page. Search, folder classification, automatic tag classification and reading mode all operate on that stored file and whatever metadata the server kept about it. This is a deliberate simplification and it has a visible consequence. There is no crawl, so there is no capture-time snapshot of a site's other pages, and no asset pass, so what survives is what the page itself carried or referenced. For an article or a documentation page that is usually the whole content. For a site that assembles itself from scripts, from lazy-loaded media or from a content delivery network, the archived copy is only as faithful as the file the extension produced, and nothing in the readme claims more. The client side is the other half of the trade, since the web client renders the stored file directly rather than re-fetching the live site.

## Conclusion

web-archive suits someone who wants their own saved pages rather than a general web archive, and the architecture makes that honest: a single file per page uploaded to storage you control, with a worker doing the indexing. Check three things first. The licence is stated twice and disagrees, so read the file in the repository before you redistribute anything. The Docker image is not a Cloudflare deployment, it is a local Workers runtime with its data in a mounted volume, which is a different set of tradeoffs from running on D1 and R2. And the image is tagged latest, so pin a version, especially since the manifest sits at 0.1.1-alpha.1 while the newest release is v0.2.0. Capture is a single HTML file, so judge it on the pages you actually want to keep rather than on the sites you can reach.

## FAQ

### How do I deploy web-archive?

Two routes. Cloudflare is marked recommended and points at a deploy document on a separate Pages site. The local route is one command, docker run -d -p 8787:8787 -v web-archive-data:/app/service/.wrangler/state with the ghcr.io/ray-d-song/web-archive image, which runs a local Workers runtime and keeps its data in that volume.

### Which licence does web-archive use?

The two sources disagree. The licence field on the repository reads GPL-3.0 while the package manifest declares ISC, so the licence file in the repository is the one to read before redistributing or embedding the code.

### What do I have to enter in the web-archive browser extension?

After deployment, the document says you enter the service address and a key. The extension saves the current page as a single HTML file and uploads it to that address, and the server stores it in the database and the storage bucket.

### Which browsers have the web-archive extension been published for?

Chrome and Firefox, with a store link for each. The workspace has a separate plugin build target for Firefox alongside the general one, and the manifest keeps the extension, the server, the shared package and the web client as separate members of one pnpm workspace.

### What does the web-archive Docker image actually run?

A local Cloudflare Workers runtime. The image is Node 24 Alpine with node-cf-worker installed globally, and it starts that binary against /app/service/wrangler.toml on port 8787, declaring /app/service/.wrangler/state as a volume where the database and bucket data live in that mode.

## Sources

- [License: GPL-3.0](https://github.com/Ray-D-Song/web-archive/blob/main/LICENSE)
- [Project website](https://web-archive-docs.pages.dev/)
- [Ray-D-Song/web-archive on GitHub](https://github.com/Ray-D-Song/web-archive)
- [README](https://github.com/Ray-D-Song/web-archive/blob/main/README.md)
- [Releases](https://github.com/Ray-D-Song/web-archive/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ray-d-song-web-archive
