Browsertrix Crawler: a browser-based web archiving crawler in one Docker container
Run a high-fidelity browser-based web archiving crawler in a single Docker container.
At a glance
- What is it?
- Browsertrix Crawler drives Brave and Chrome through Puppeteer to capture high-fidelity page archives as WACZ files. It suits archivists and engineers who need real browser rendering, and it is heavy machinery for anyone who just wants a list of links.
- Who is it for?
- Adopt Browsertrix Crawler if you need rendered, high-fidelity captures and can give a container the NET_ADMIN and SYS_ADMIN capabilities plus a 1gb shared memory segment. Do not adopt it for a quick link inventory or for a host where you cannot run Docker with those privileges.
- Can I use it commercially?
- Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
- Is it still maintained?
- Yes. The repository last received commits 7 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap Browsertrix Crawler fills for web archiving
A conventional crawler fetches HTML and follows links. That is enough for indexing text, and it is not enough for archiving a page as a person saw it. Modern sites assemble themselves in the browser: scripts fetch data, fonts and images load late, and the DOM that arrives from the server is a skeleton. Browsertrix Crawler exists for that second problem. The README describes it as a "standalone browser-based high-fidelity crawling system", and the phrase high-fidelity is doing the work. It runs one or more Brave Browser windows in parallel through Puppeteer and records what the browser actually does, using the Chrome Devtools Protocol. The output is an archive of the rendered page, not just the response body. The audience is narrow on purpose: digital preservation teams, national and institutional archivists, and engineers building archiving pipelines. The project's own history points the same way. Initial functionality was developed for the zimit project in a collaboration between Webrecorder and Kiwix, with support from Portico for the 0.4.x line. This is infrastructure for people who are accountable for what a page contained on a given day.
How the crawl actually works: Puppeteer, Brave and CDP
The mechanism is worth understanding before you run anything, because it explains the resource requirements. Browsertrix Crawler does not implement its own HTTP client for page capture. It launches browser instances and controls them with Puppeteer, then captures data through CDP. The Dockerfile shows that the container is built on a browser base image, webrecorder/browsertrix-browser-base, pinned by a BROWSER_VERSION build argument (1.93.138 in the file), and it sets BROWSER_BIN=google-chrome. The image exposes ports 9222 and 9223 for the browser debugging interfaces and 6080 for a VNC endpoint, which is how you can watch a crawl in progress. Dependencies in package.json fill in the rest of the pipeline: puppeteer-core for control, warcio for WARC-format records, @webrecorder/wabac for capture and replay support, browsertrix-behaviors for page-level automation, and client-zip for producing the archive bundle. The crawl is coordinated through a queue (p-queue) with ioredis present for queue state, and robots-parser enforces robots.txt handling. A single container therefore holds the browser, the crawler logic and the archive writer. That is the whole design bet: one image, no separate browser farm to operate.
Installing Browsertrix Crawler with Docker and running a first crawl
The README points to the hosted documentation at crawler.docs.browsertrix.com for usage, and the repository ships both a Dockerfile and a docker-compose.yml. The compose file is the shortest path to a working setup, because it already declares the privileges and shared memory the browser needs.
The service uses the image webrecorder/browsertrix-crawler:latest, mounts ./crawls on the host to /crawls in the container, adds the NET_ADMIN and SYS_ADMIN capabilities, and sets shm_size to 1gb. That shared memory setting is not decorative; Chromium-based browsers crash or stall without adequate /dev/shm. The REGISTRY variable is left for you to override if you pull from a mirror.
services:
crawler:
image: ${REGISTRY}webrecorder/browsertrix-crawler:latest
volumes:
- ./crawls:/crawls
cap_add:
- NET_ADMIN
- SYS_ADMIN
shm_size: 1gbStart it with docker compose up, and the container's entrypoint (docker-entrypoint.sh in the repository root) runs the crawler. Output archives land under the mounted /crawls directory on your host. If you prefer a plain docker run, you must reproduce the same volume and capability flags yourself; the compose file is the reference for what is required.
The image is built from a browser base and sets a VNC password via the VNC_PASS environment variable (vncpassw0rd! in the Dockerfile) and a display geometry of 1360x1020x16. Port 6080 is exposed for that VNC session, so you can attach a viewer to port 6080 and watch pages load while the crawl runs. The documentation at crawler.docs.browsertrix.com is where the crawl configuration options live; the README itself does not enumerate them, so treat the docs site as the source for seed URLs, scope rules and limits.
Where Browsertrix Crawler is the wrong tool
The container needs NET_ADMIN and SYS_ADMIN capabilities and a 1gb shared memory segment. On a managed container platform that forbids added capabilities, or on a host where you cannot raise shm_size, the crawl will not run as documented. That rules out a lot of serverless and locked-down Kubernetes environments unless you control the pod security policy. The second constraint is cost per page. Launching a real browser, rendering scripts and capturing CDP traffic is orders of magnitude heavier than an HTTP fetch. For a site of static HTML pages, you are paying browser overhead for nothing. Third, the project is version 1.15.0-beta.0 at the most recent tag, with 1.14.3 and 1.14.2 as the stable releases before it. If you need a predictable artifact, pin a stable tag rather than latest. Finally, the README is thin by design: it delegates usage to the hosted docs and does not document rollback, failure recovery or resume behaviour. If your workflow depends on restarting an interrupted crawl cleanly, confirm that in the docs before you build a pipeline around it.
How it differs from wget-style and WARC-only crawlers
The obvious alternative for many teams is a command-line fetcher that writes WARC output, or a headless crawler that renders pages but does not archive them. The difference is where the capture happens. A fetcher-based crawler records the HTTP exchange: request, response, headers, body. That is a faithful record of the server's answer and a poor record of the page, because anything assembled by JavaScript is missing or captured in its pre-render state. Browsertrix Crawler inverts this. The browser is the primary client, and CDP is the capture channel, so what gets recorded is what the rendering engine produced, including resources fetched after the initial document. The trade-off is the one described above: you inherit browser resource demands, capability requirements and startup cost. There is also a packaging difference. The project depends on warcio for WARC records and client-zip for bundling, and its Python requirements file pins wacz>=0.5.0, so the intended deliverable is a WACZ archive rather than a loose set of WARC files. If your downstream system expects WARC and nothing else, check whether the WACZ output can be unpacked to what you need before adopting.
Maintenance, releases and the AGPL-3.0 licence
The repository is not archived, and the last push was on 2026-08-27, which coincides with the v1.15.0-beta.0 tag. Release cadence in the visible window is tight: 1.14.2 on 2026-08-13, 1.14.3 on 2026-08-20, and the beta a week later. That pattern suggests active work, and it also means the beta tag is not the one to pin in production. The upgrade cost sits mostly in the browser base image. The Dockerfile takes BROWSER_VERSION as a build argument and derives the base image from it, so a browser bump changes the rendering engine underneath your captures. If you care about reproducing an archive byte for byte, record which image tag produced it. On licensing, package.json declares AGPL-3.0-or-later and the README points to the LICENSE file for details. The AGPL's network clause matters if you modify the crawler and expose it as a service to others; that is a question for your own counsel, not something this article can settle. Note also that the Dockerfile downloads an ad host blocklist at build time from a GitHub raw URL, so an offline or air-gapped build needs that step accounted for.
Editorial conclusion
Adopt Browsertrix Crawler if you need rendered, high-fidelity captures and can give a container the NET_ADMIN and SYS_ADMIN capabilities plus a 1gb shared memory segment. Do not adopt it for a quick link inventory or for a host where you cannot run Docker with those privileges. Before committing, verify the image tag you intend to pin, confirm the /crawls volume path matches your storage layout, and read the AGPL-3.0 licence terms if you plan to redistribute a modified version.
Frequently asked questions
How do I install Browsertrix Crawler?
The repository ships a Dockerfile and a docker-compose.yml; the compose file uses the image webrecorder/browsertrix-crawler:latest, mounts ./crawls to /crawls, adds NET_ADMIN and SYS_ADMIN capabilities and sets shm_size to 1gb. Running docker compose up starts the crawler through its entrypoint script. The README directs readers to crawler.docs.browsertrix.com for full usage details.
Does Browsertrix Crawler work on Windows or Ubuntu?
The project is distributed as a Docker container built on a Linux browser base image, so the practical requirement is a host that can run that container with the NET_ADMIN and SYS_ADMIN capabilities and a 1gb shared memory segment. The README does not list supported host operating systems. Docker Desktop on Windows and a Linux host such as Ubuntu can both run containers, but the repository does not document either one specifically.
What does the Browsertrix Crawler output look like?
The dependencies point to WARC-format records through warcio and to client-zip for bundling, and requirements.txt pins wacz>=0.5.0, so the deliverable is a WACZ archive. Archives are written under the container path /crawls, which the compose file maps to ./crawls on the host. The README does not describe the internal file layout of the bundle.
Can I watch a crawl while it is running?
The Dockerfile exposes ports 9222, 9223 and 6080, and sets a VNC password through the VNC_PASS environment variable with a display geometry of 1360x1020x16. Port 6080 is the VNC endpoint, so a viewer attached there can show the browser session during a crawl. The README does not document the VNC workflow itself.
Is Browsertrix Crawler the same as Browsertrix Cloud?
The repository describes only the standalone crawler, a system designed to run in a single Docker container, and the README does not mention a hosted or cloud product. Nothing in the repository documentation describes a cloud service, an account or a login. Treat the crawler and any hosted offering as separate subjects until you confirm it in the official documentation.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/webrecorder-browsertrix-crawler)