webclaw: local-first web extraction in Rust, with an MCP server for agents
Fast, local-first web content extraction for LLMs. Scrape, crawl, extract structured data — all from Rust. CLI, REST API, and MCP server.
At a glance
- What is it?
- webclaw turns URLs into markdown, JSON or LLM-ready text from a Rust CLI, an MCP server or a self-hosted REST API. The local path handles ordinary pages; bot-protected and JavaScript-heavy ones need the hosted service or your own key.
- Who is it for?
- Adopt webclaw if you want extraction to run on your own machine, drive it from an MCP client, and keep the AGPL-3.0 obligations in mind before you ship a modified server. Skip it if your targets are JavaScript-rendered or bot-protected and you do not intend to run a browser, a proxy pool or the hosted API.
- Can I use it commercially?
- Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
- Is it still maintained?
- Yes. The repository last received commits 7 days ago.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The two bad outputs webclaw is built to avoid
Scrapers tend to hand an agent one of two things. Either a blocked page, a login wall or an empty app shell, or raw HTML with nav, scripts, styling, ads and duplicated boilerplate still attached. The README frames the problem exactly that way, and the fix it proposes is a single conversion step: a URL in, clean content out.
The intended audience is narrow and identifiable. People building RAG indexes, agent tool calls and content pipelines who want the extraction to happen on their own machine. The README's own example is the shortest statement of scope:
webclaw https://example.com --format markdownWhat comes back is a markdown document with the page's heading and body text, not a DOM dump. That is the whole pitch. It is not a browser automation framework, and it does not try to be one.
How the Rust workspace splits fetching, extraction and LLM work
The repository is a Cargo workspace with members under crates/, and the workspace dependencies name the pieces: webclaw-core, webclaw-fetch, webclaw-llm and webclaw-pdf, plus binaries for the CLI, MCP server and server. That layout tells you the fetch layer is separable from the extraction core, and the LLM-dependent features sit in their own crate rather than being wired into every command.
The Dockerfile is more explicit about what ships. It builds three binaries: webclaw for single-shot extraction and crawl, webclaw-mcp as a stdio MCP server for agents, and webclaw-server as a minimal REST API for self-hosting. Its header comment is unusually blunt: this is not the hosted API at api.webclaw.io, because the cloud service adds anti-bot bypass, JS rendering, multi-tenant auth and async jobs that are intentionally not open-source. Treat that as the architectural boundary. Everything local is a plain fetch plus parsing pipeline; the hard anti-bot work lives on the other side of a paid service.
The tool table in the README marks scrape, crawl, map, batch, extract, summarize and diff as local, with extract and summarize noted as "yes, with local or configured LLM". That phrasing matters: those two commands are only as good as the model you point them at, and the docker-compose file shows the intended local arrangement.
Installing webclaw and running a first extraction
The README lists five install paths: an npx installer for MCP clients, Homebrew, prebuilt binaries from GitHub Releases, Docker, and cargo install from git. The fastest for an agent setup is the installer, which detects supported clients and writes their MCP configuration:
npx create-webclawIf you would rather not run an installer, point any MCP client at the npx launcher directly. This is the configuration block the README gives, and it needs no local install at all:
{
"mcpServers": {
"webclaw": {
"command": "npx",
"args": ["-y", "@webclaw/mcp"]
}
}
}For a plain terminal check, the Docker route avoids a Rust toolchain entirely:
docker run --rm ghcr.io/0xmassi/webclaw https://example.comYou should see the extracted content printed to stdout. From there the flags are the interesting part. Restricting to the article body and stripping chrome is a single invocation:
webclaw https://example.com/blog/post --only-main-content --exclude "nav, footer, .sidebar, .ad"For a documentation site, crawl mode follows same-origin links with a depth and page cap:
webclaw https://docs.rust-lang.org --crawl --depth 2 --max-pages 50Building from source needs native tools. The README's prerequisites table lists pkg-config, libssl-dev, cmake, clang and build-essential on Debian and Ubuntu, the openssl-devel equivalents on Fedora, and xcode-select on macOS. The Dockerfile confirms cmake and clang are there for BoringSSL, which is the TLS stack behind wreq. If a source build fails, that is the first thing to check.
Self-hosting with docker-compose and a local Ollama model
The compose file defines two services. webclaw builds from the local Dockerfile and publishes port 3000 by default, overridable through WEBCLAW_PORT. ollama runs the official image with a named volume and is set as a dependency, and webclaw receives OLLAMA_HOST=http://ollama:11434 in its environment. That is how the extract and summarize commands get a model without sending page content to a third party.
The compose file ships CPU-only and leaves the GPU reservation commented out, so the default path is slow for anything but small models. It also notes that a model has to be pulled after startup:
docker compose exec ollama ollama pull qwen3:1.7bThe healthcheck is worth reading before you trust a deployment. It runs webclaw --help every 30 seconds with a 5 second timeout. That confirms the binary starts. It does not confirm that fetching works, that DNS resolves inside the container, or that the Ollama dependency is reachable. If you wire this into anything that pages a human, the healthcheck will stay green while extraction is broken.
The README also mentions setting WEBCLAW_API_KEY to handle bot-protected and JavaScript-rendered pages, which routes those requests to the hosted service rather than solving them locally.
Where webclaw stops: no browser, no anti-bot bypass, AGPL-3.0
The clearest limitation is stated by the project itself. The Dockerfile says the cloud service adds anti-bot bypass and JS rendering that are intentionally not open-source. So if your targets are single-page apps, Cloudflare-fronted sites or pages behind a challenge, the local binary is the wrong tool unless you supply your own proxy and rendering layer. The repository does include proxies.example.txt and an examples/proxy-backed-crawling directory, which suggests proxies are a supported configuration, but the README does not document a bundled browser engine.
Licensing is the second constraint. The workspace package declares AGPL-3.0, and the repository carries a LICENSE file at the root. For internal tooling that nobody outside your organisation interacts with, that is usually unremarkable. For a modified webclaw-server offered to users over a network, the AGPL's source-availability expectations are the thing to read with your own counsel. Nothing here is legal advice, and the licence text is the only authority.
The healthcheck gap described above is a third, smaller failure mode: a container that reports healthy while every fetch fails.
Maintenance is not a concern on the evidence available. The last push was on 2026-09-09, and the most recent release listed is v0.6.22 from 2026-08-30, with v0.6.21 and v0.6.20 both on 2026-08-16. The workspace version in Cargo.toml reads 0.6.23, one ahead of the newest release, which is normal for a main branch between tags. Upgrade cost is low for the CLI: the interface is flags on a single command, and the changelog at the repository root is where breaking changes would appear.
How webclaw differs from Firecrawl and Crawl4AI
The repository's own topics list it alongside firecrawl-alternative, crawl4ai-alternative, jina-alternative, apify-alternative, scraperapi-alternative and scrapingbee-alternative, and there is an examples/firecrawl-compatible-api directory. So the project is positioning itself in a crowded field rather than pretending the field is empty.
The real difference is where the browser lives. Firecrawl and Crawl4AI are built around rendering pages in a headless browser; that is what lets them return content from JavaScript-driven sites, and it is also why they are heavier to run and harder to keep unblocked. webclaw's local path is a Rust fetch and parse pipeline with no browser in the loop, which is why the Docker image stays small enough to run a CLI in one container and why a scrape returns in the time a request takes. The trade is coverage: the moment a page needs rendering or a challenge solved, the local binary is out and you are back to the hosted API, a proxy, or a different tool.
If your corpus is documentation, blogs, marketing pages and server-rendered content, that trade is a good one. If it is dashboards and app shells, it is not.
What the README does not answer
Several things a team would want before committing are simply absent. There is no documented rollback procedure for a server upgrade, no stated rate limit for the self-hosted REST API, and no description of how crawl concurrency is bounded beyond the --max-pages flag. The README also does not document what happens when --include and --exclude selectors overlap, which is the kind of detail that decides whether a pipeline is reproducible.
The benchmarks/ and targets_1000.txt entries at the repository root suggest measurement work exists, but nothing in the README states a throughput or success-rate figure, and none should be inferred from directory names. If extraction quality on your specific corpus is the deciding factor, the only honest test is running the CLI against a sample of your own URLs and reading the output.
Editorial conclusion
Adopt webclaw if you want extraction to run on your own machine, drive it from an MCP client, and keep the AGPL-3.0 obligations in mind before you ship a modified server. Skip it if your targets are JavaScript-rendered or bot-protected and you do not intend to run a browser, a proxy pool or the hosted API. Verify first that the OSS Docker image covers your pages: the Dockerfile states the hosted service adds anti-bot bypass, JS rendering, multi-tenant auth and async jobs that are intentionally not open-source, so check the self-hosting docs before you promise coverage to anyone.
Frequently asked questions
What is the best web crawler for AI?
There is no single answer, and webclaw's own README frames the choice as a trade between blocked or empty pages and raw HTML full of boilerplate. webclaw's local path handles ordinary server-rendered pages with no API key, while bot-protected and JavaScript-rendered pages require the hosted service or your own proxy and rendering setup.
How do I install webclaw?
The README lists five routes: npx create-webclaw for MCP clients, brew tap 0xMassi/webclaw followed by brew install webclaw, prebuilt binaries from GitHub Releases, the ghcr.io/0xmassi/webclaw Docker image, and cargo install from the git repository. Source builds need native tools such as pkg-config, libssl-dev, cmake and clang.
Does webclaw need an API key?
Most sites extract locally with no API key, according to the README. Setting WEBCLAW_API_KEY is what the documentation gives as the way to handle bot-protected and JavaScript-rendered pages.
Can I self-host webclaw?
Yes. The Dockerfile builds a webclaw-server binary described as a minimal REST API for self-hosting, and docker-compose.yml wires it to an Ollama service on port 3000 by default. The Dockerfile states this is not the hosted API at api.webclaw.io, because anti-bot bypass, JS rendering, multi-tenant auth and async jobs are not open-source.
What licence does webclaw use?
The workspace package in Cargo.toml declares AGPL-3.0, and the repository carries a LICENSE file at the root. Anyone modifying the server and exposing it over a network should read that licence text directly.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/0xmassi-webclaw)