# fastCRW: a self-hosted Rust scraper, crawler and search API with a Firecrawl-compatible surface

> fastCRW (the us/crw repository) bundles scrape, crawl, map, search and extract behind one Rust binary and an MCP server for AI agents. The interesting part is the Firecrawl-shaped API and the AGPL split between engine and SDKs, not the benchmark claims.

**us/crw** — Fast, lightweight Firecrawl/Tavily alternative in Rust. Web scraper, crawler & search API with MCP server for AI agents. Drop-in Firecrawl-compatible API (/scrape, /crawl, /search). 2.3x faster than Tavily, 1.5x faster than Firecrawl in 1K-URL benchmarks. 6 MB RAM, single binary. Self-host or use managed cloud.

- Repository: https://github.com/us/crw
- Website: https://fastcrw.com
- Stars: 1,091 · Forks: 88
- Language: Rust
- License: AGPL-3.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/us-crw

## The problem fastCRW targets: five scraping operations behind one Firecrawl-shaped API

Most teams that feed web pages to a model end up with three separate dependencies: something that turns a URL into markdown, something that walks a site, and something that searches. fastCRW collapses those into one engine. The README lists five operations: scrape, crawl, map, search and extract. Scrape returns markdown, HTML, links, screenshots or schema JSON from a single URL. Crawl follows a bounded site crawl and collects its pages. Map discovers URLs without scraping every page. Search queries the web and can optionally scrape selected results. Extract produces structured fields from one or many URLs.

The audience is narrower than the tagline suggests. This is for engineers who already call Firecrawl-style endpoints and want the same routes on their own hardware, or who are wiring an MCP server into Claude Code, Cursor, Codex, Gemini CLI, OpenCode or Windsurf. It is not a general-purpose crawler for non-programmers, and the README offers no GUI.

The compatibility claim is the load-bearing one. The repository carries both COMPATIBILITY.md and COMPATIBILITY-firecrawl.md, and the description names /scrape, /crawl and /search as drop-in routes. That is the actual selling point: migration cost measured in request bodies, not in a rewrite.

## How the Rust workspace is split, and where SearXNG fits

Cargo.toml defines a workspace with eleven members: crw-core, crw-diff, crw-renderer, crw-extract, crw-crawl, crw-search, crw-server, crw-mcp, crw-mcp-proto, crw-browse and crw-cli. That layout tells you the architecture better than the README does. The CLI, the HTTP server and the MCP server are three separate binaries over shared crates, so the same crawl and extract logic serves a terminal call, a REST request and an agent tool call.

Rendering is pluggable rather than baked in. The workspace comments describe an optional Chrome TLS/JA3/HTTP2 impersonation client, referenced only by crw-renderer behind an `impersonated` feature, and note that wreq-util 0.2 is Apache-2.0 while the later 3.0.0-rc line is LGPL and deliberately avoided. The Dockerfile comments describe a build that defaults to shipping crw, crw-server and crw-mcp, with production overriding the package set so the engine container only runs crw-server.

docker-compose.yml is where the real dependencies appear. The crw service depends on `lightpanda` as a started service, on an optional `chrome` profile, on an optional `chrome-stealth` profile, and on `searxng` with `condition: service_healthy`. That last one matters: the compose comments state the SearXNG sidecar backs `/v1/search`. A self-hosted fastCRW search endpoint is therefore not a standalone capability. It is a proxy in front of a search engine you also have to run. The compose file also notes the stealth profile uses a browserless image under SSPL and should only be run if you accept those terms, and that the heavy and stealth profiles are mutually exclusive in practice.

One detail worth reading in full is the DNS comment in Cargo.toml. The resolver deliberately keeps reqwest's default getaddrinfo path instead of hickory-dns, because getaddrinfo reads /etc/resolv.conf and resolves Docker's embedded DNS at 127.0.0.11, which the hickory resolver failed on for direct egress. It runs on tokio's blocking pool. That is the kind of note that only exists after someone lost an afternoon to it.

## Installing fastCRW and running a first scrape

The README gives a one-command install for macOS and Linux, Intel and ARM. It runs locally and free with no account:

```bash
curl -fsSL https://fastcrw.com/install | sh
```

After that, the CLI is available as `crw`. The README's first example is a scrape, which it says works right after install:

```bash
crw https://example.com
```

Search is a second step. The README states search works after `crw setup`, which is also the interactive path for choosing which AI tools to register the MCP server with:

```bash
crw search "rust tutorials"
```

If you have a managed key, the same install command takes it as an environment variable and additionally registers the MCP server with detected AI coding tools. The README says the key is written to `~/.config/crw/config.toml` rather than copied into each tool, and that `CRW_NO_AGENTS=1` skips the agent registration step:

```bash
curl -fsSL https://fastcrw.com/install | CRW_API_KEY=crw_live_... sh
```

For the MCP path without the install script, the README gives a single npx command:

```bash
npx -y crw-mcp@latest install
```

From Python, the SDK is a separate package. The README exports the key once, installs `crw`, and calls `scrape` with a formats list, then reads `page["markdown"]`:

```python
from crw import CrwClient

client = CrwClient()
page = client.scrape("https://example.com", formats=["markdown"])

print(page["markdown"])
```

For self-hosting the server rather than the CLI, the repository ships docker-compose.yml with the container port fixed at 3000. The compose comments state the host bind address and port are overridden through `CRW_BIND_ADDRESS` and `CRW_HOST_PORT`, that these are Compose-only host-side interpolation and are not passed into the container, and that setting `CRW_BIND_ADDRESS=127.0.0.1` stops publishing on public interfaces when you front it with a reverse proxy:

```yaml
services:
  crw:
    ports:
      - "${CRW_BIND_ADDRESS:-0.0.0.0}:${CRW_HOST_PORT:-3000}:3000"
```

If you build from source instead, the README states the workspace requires Rust 1.85 or newer and shows `make check-fast` after cloning.

## Where fastCRW stops being the right tool

The benchmark numbers are the weakest part of the pitch, and the README presents two different ones. The first is a comparison on Firecrawl's own public 1,000-URL dataset, where the README says fastCRW recovered more truth than Crawl4AI and Firecrawl and idled at roughly 14 MB RAM. The second is answer accuracy on 600 AA-Omniscience questions, where the README claims 90.0 percent. The repository description quotes different figures again: 2.3x faster than Tavily and 1.5x faster than Firecrawl, at 6 MB RAM. Those numbers do not agree with each other, and the README itself points at BENCHMARKS.md for methodology and reproduction. Treat any of them as a vendor measurement until you reproduce them on your own URLs. Nothing in the repository establishes how these translate to a corpus that is mostly JavaScript-heavy or behind a login.

Rendering is the second boundary. The base compose stack depends on lightpanda; full browser rendering is opt-in through the `chrome` profile, and stealth rendering through `chrome-stealth`, which the compose comments flag as an SSPL-licensed browserless image. If your targets need real JS execution, you are not running the small single-binary configuration described in the tagline. You are running a Chrome sidecar, and in the stealth case you are accepting an additional licence.

Search has a hard dependency too. Because `/v1/search` is backed by the SearXNG sidecar, a self-hosted deployment that skips SearXNG has no search. The compose file waits for that service to be healthy precisely so the first request does not race the cold start. If you wanted a search API without operating a search engine, the managed deployment is the only path the README describes.

Finally, the README states plainly that capabilities and response shapes can differ by deployment, pointing at a `/v1/capabilities` endpoint and a response-shapes page. A drop-in replacement that returns a different shape depending on whether it is managed or local is a compatibility risk you have to test, not assume.

## fastCRW against Firecrawl and Tavily: same routes, different operating model

The honest comparison is not speed. It is who runs the browser. Firecrawl and Tavily are hosted services: you send a request, they own the proxies, the renderers and the scaling, and you pay per call. fastCRW offers the same shape of API but moves the runtime to your machine. The README's own table makes this explicit, listing managed API as best for zero infrastructure and managed scaling, and local or self-hosted as best for data control, private networks or custom infrastructure.

That difference has concrete consequences. With a hosted service, a target site blocking your crawler is the vendor's problem. With self-hosted fastCRW, the compose comments show you choosing renderers, search, auth, proxies and capacity yourself, including whether to run the heavier Chrome profiles. Your egress IPs are your own, which is better for private networks and worse for sites that rate-limit by IP range.

The compatibility layer is what makes the swap plausible. Because the routes mirror Firecrawl's, a migration is largely a base-URL change plus whatever the COMPATIBILITY-firecrawl.md page says does not line up. That is a much smaller project than adopting a crawler with its own request schema. The cost moves from per-request billing to operating a container, a search sidecar and possibly a browser.

Crawl4AI appears in the README's benchmark comparison but is not described in enough detail there to compare architectures. If you are choosing between the two, BENCHMARKS.md is the document to read, and the reproduction instructions in it are the thing to actually run.

## Licence and upgrade cost: AGPL-3.0 engine, MIT SDKs, release cadence

The licence split is stated in the README and repeated in Cargo.toml. The engine and the MCP server are AGPL-3.0. The Python and TypeScript SDKs are MIT. The embedding licence is a separate arrangement handled by email. The workspace package metadata also carries AGPL-3.0.

AGPL-3.0 is the decision that will filter adopters faster than any performance number. If you run a modified crw-server as a network service that other people interact with, the licence's network clause is the thing your legal review will want to look at, and the repository does not offer guidance on it beyond naming the licence and the separate embedding contact. I am not giving legal advice here; the point is that a permissive-licence assumption is wrong for this project, and the SDKs being MIT does not change what the server binary is under.

The upgrade cost is visible in the release history. Three tagged releases appear in the recent list, v0.32.0 on 2026-08-24, v0.33.0 on 2026-09-02 and v0.34.0 on 2026-09-08, while Cargo.toml declares workspace version 0.35.1. That is a fast-moving pre-1.0 line with frequent minor bumps. The repository has release-please configuration and a CHANGELOG.md, so version-to-version changes are documented, but nothing in the repository promises API stability across those bumps. The last push to the default branch was on 2026-09-09, and the repository is not archived.

If you self-host, budget for the sidecars rather than the binary. The Dockerfile comments describe a cargo-chef dependency layer and a production flow that bakes the expensive external-crate compile as a tagged image so the nightly docker prune cannot evict it. That is a real operational pattern you inherit, along with the lightpanda and searxng services, if you deploy the compose stack as written.

## Conclusion

Adopt fastCRW if you want a small self-hosted engine that speaks the Firecrawl routes your code already calls, and if AGPL-3.0 on the engine is acceptable for how you deploy it. Do not adopt it if you need the managed proxies, JS rendering and hosted billing to come from someone else while keeping the licence permissive, or if you cannot run the SearXNG sidecar that backs /v1/search. Before committing, verify three things against your own traffic: that the Docker Compose stack comes up with the renderer profile you intend to use, that /v1/capabilities on your build lists the operations you actually call, and that your Firecrawl request bodies survive the migration page's mapping. The response shapes can differ between the managed API and a local build, so pin one deployment and test against it.

## FAQ

### How do I install fastCRW?

The README gives a one-command install for macOS and Linux on Intel and ARM: piping https://fastcrw.com/install to sh. The same command accepts CRW_API_KEY as an environment variable to connect a managed key, and CRW_NO_AGENTS=1 to skip registering the MCP server with detected AI tools.

### Does fastCRW have an MCP server for AI agents?

Yes. The README documents an MCP server installed either through the install script, through crw setup, or directly with npx -y crw-mcp@latest install. It states that Claude Code, Cursor, Codex, Gemini CLI, OpenCode and Windsurf are detected automatically when already set up.

### Is fastCRW a drop-in replacement for Firecrawl?

The project describes a Firecrawl-compatible API covering /scrape, /crawl and /search, and the repository includes COMPATIBILITY-firecrawl.md and a migration page. The README also warns that capabilities and response shapes can differ by deployment, so the compatibility should be verified against your own requests.

### What licence is fastCRW under?

The README states the engine and MCP server are AGPL-3.0, while the Python and TypeScript SDKs are MIT. Cargo.toml declares AGPL-3.0 for the workspace, and the README lists a separate embedding licence contact.

### Can I self-host the fastCRW search endpoint?

The docker-compose.yml file lists a SearXNG sidecar that the comments say backs /v1/search, and the crw service waits for it to report healthy. A self-hosted deployment therefore needs that search service running alongside the engine.

### Does fastCRW need a browser to scrape pages?

The base compose stack depends on lightpanda, and the comments describe chrome and chrome-stealth as opt-in profiles, with the stealth profile using an SSPL-licensed browserless image. Full browser rendering is therefore a configuration choice rather than the default.

## Sources

- [License: AGPL-3.0](https://github.com/us/crw/blob/main/LICENSE)
- [Project website](https://fastcrw.com)
- [README](https://github.com/us/crw/blob/main/README.md)
- [Releases](https://github.com/us/crw/releases)
- [us/crw on GitHub](https://github.com/us/crw)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/us-crw
