Model or dataset
firecrawl/firecrawl avatar
firecrawl/firecrawl

Firecrawl: a self-hostable web context API for agents

The API to search, scrape, and interact with the web at scale. 🔥

186,493 stars9,982 forksTypeScriptAGPL-3.0

At a glance

What is it?
Firecrawl turns URLs into Markdown or structured JSON through search, scrape, interact, crawl and map endpoints. It ships a hosted service, a Docker Compose self-host path and an AGPL-3.0 core, and the licence is the part most teams underestimate.
Who is it for?
Adopt Firecrawl if you need Markdown or structured JSON out of JS-heavy pages and you are willing to either pay the hosted service or run the API, Redis, Postgres and a Playwright microservice yourself. Do not adopt it if you only need static HTML parsing, since a plain HTTP client plus an HTML parser avoids the whole stack.
Can I use it commercially?
Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

The problem Firecrawl solves, and who it is actually for

Fetching a page is easy. Getting usable text out of it is not. The README frames the target as agents and LLM applications that need "clean Markdown or structured data your agents can ship with", and the feature list backs that up: rotating proxies, JavaScript rendering, rate limits and media parsing are described as handled by the service rather than by your code. The README claims coverage of "96% of the web, including JS-heavy pages" and a P95 latency of 3.4s, both cited to a benchmark post on the project's own blog rather than to an independent test.

That framing tells you who this is for. If you are building a retrieval pipeline, a research agent, a lead enrichment job or a newsletter generator, the interesting work is downstream of the fetch, and Firecrawl is trying to sell you the fetch as a managed primitive. The repository's examples directory confirms the audience: R1_company_researcher, deepseek-v3-company-researcher, crm_lead_enrichment, deep-research-apartment-finder. These are LLM application patterns, not general web scraping.

If you are doing high-volume, stable, well-structured scraping of sites you control or already understand, Firecrawl is the wrong layer. You would be paying for browser orchestration and proxy rotation you do not need. The project is also not a scraping framework in the Scrapy sense: there is no spider class, no middleware chain, no selector DSL. It is an HTTP API with SDKs wrapped around it.

How Firecrawl works: five endpoints over a queue, a browser pool and Postgres

The public surface is small and the README names it directly. Search returns web results with full page content. Scrape converts one URL to markdown, HTML, screenshots or structured JSON. Interact operates on a page after scraping it. Crawl walks all URLs of a site from a single request. Map discovers URLs. Batch Scrape handles thousands of URLs asynchronously. The README groups Search, Scrape and Interact as core endpoints and Agent, Crawl, Map and Batch Scrape as the wider set.

The interesting part is Interact, because it forces a stateful design. You scrape a URL, the response carries a scrape_id in its metadata, and you pass that id to a separate interact call with a prompt. The documented example scrapes amazon.com, then sends "Search for 'mechanical keyboard'" and "Click the first result" as two sequential prompts. The response includes an output field and a liveViewUrl pointing at liveview.firecrawl.dev. That means the service holds a live browser session keyed by scrape id, which is why self-hosting is not a single container.

The docker-compose.yaml confirms the shape. The API service depends on REDIS_URL for queuing and rate limiting (REDIS_RATE_LIMIT_URL defaults to the same instance), a Postgres instance addressed as nuq-postgres, and PLAYWRIGHT_MICROSERVICE_URL defaulting to http://playwright-service:3000/scrape. Concurrency is bounded by explicit environment variables: NUM_WORKERS_PER_QUEUE defaults to 8, CRAWL_CONCURRENT_REQUESTS to 10, MAX_CONCURRENT_JOBS to 5, and BROWSER_POOL_SIZE to 5. Those four numbers are the real capacity model, and they are conservative by default. A self-hosted instance that feels slow is usually a browser pool of five, not a bug.

The compose file also lists OPENAI_API_KEY, OPENAI_BASE_URL, MODEL_NAME, MODEL_EMBEDDING_NAME and OLLAMA_BASE_URL, which tells you the AI-backed features expect an LLM endpoint you supply when self-hosting. USE_DB_AUTHENTICATION defaults to false, so a local stack starts without an auth database unless you change it.

Installing Firecrawl and running a first scrape

There are two routes. The hosted route needs only an API key from firecrawl.dev, and the README's Quick Start uses it throughout. The Python SDK is installed as firecrawl and constructed with your key.

python
from firecrawl import Firecrawl

app = Firecrawl(api_key="fc-YOUR_API_KEY")

result = app.scrape('firecrawl.dev')

The Node SDK is imported as firecrawl and takes an apiKey option. If you would rather not install an SDK, the REST endpoint is POST https://api.firecrawl.dev/v2/scrape with an Authorization: Bearer header and a JSON body containing the url field.

bash
curl -X POST 'https://api.firecrawl.dev/v2/scrape' \
-H 'Authorization: Bearer fc-YOUR_API_KEY' \
-H 'Content-Type: application/json' \
-d '{
  "url": "firecrawl.dev"
}'

What you should see is Markdown with a top-level heading and a Features list, not raw HTML. The README's sample output begins with "# Firecrawl" and a bulleted feature list. If you get a wall of div tags, you are hitting the site rather than the API.

There is also a CLI, and the README shows it scraping and then trimming to the main content.

bash
firecrawl scrape https://firecrawl.dev
firecrawl https://firecrawl.dev --only-main-content

For the self-hosted route, the repository ships docker-compose.yaml at the top level and a SELF_HOST.md file. The compose file builds from apps/api by default, with a commented-out alternative pointing at the ghcr.io/firecrawl/firecrawl image. If you want the published image instead of a local build, that is the line to swap. The service definition raises the file descriptor limit to 65535 soft and hard, which is a hint about how many concurrent sockets a browser-driven crawler opens.

To wire Firecrawl into an MCP client, the README gives a JSON block under the mcpServers key with a firecrawl-mcp entry whose command is npx and whose args begin with -y firecrawl-mcp. The README also documents a one-command skill install for coding agents.

bash
npx -y firecrawl-cli@latest init --all --browser

The README states you should restart your agent after installing, and names Claude Code, Antigravity and OpenCode as supported.

Where Firecrawl breaks down: state, cost and the self-host ceiling

The Interact endpoint is the clearest limitation. It is stateful by design: you scrape, you get a scrape_id, you issue prompts against that id. Anything that scales horizontally has to route subsequent interact calls back to the instance holding that browser session, or accept that the session may be gone. The README does not document session lifetime, expiry, or what happens when a scrape_id goes stale. That is a real gap for anyone building a long-running agent loop.

The second constraint is that the headline reliability numbers are self-reported. The README links its 96% coverage and 3.4s P95 figures to a post on firecrawl.dev. Those may be accurate, but they are vendor benchmarks, and the README does not describe the methodology. Treat them as marketing until you reproduce them on your own targets.

Third, self-hosting is not a single binary. You are running the API, Redis for queueing and rate limiting, Postgres, and a Playwright microservice, and you are supplying an LLM endpoint if you want the AI-backed features. The AGENTS.md, CLAUDE.md and CONTRIBUTING.md files at the repository root suggest the project is developed with agent tooling in the loop, which is fine, but it does not reduce the operational surface. A team without Docker Compose experience will spend their first day on infrastructure rather than on scraping.

Finally, the wrong-tool case is worth stating plainly. If your target pages are server-rendered and stable, an HTTP client plus an HTML parser will be faster, cheaper and easier to debug than a browser pool. Firecrawl earns its place when JavaScript execution, proxy rotation or anti-bot handling is the actual problem.

Firecrawl alternatives and the real difference in approach

The closest comparison in spirit is a general-purpose scraping framework such as Scrapy. The difference is architectural, not cosmetic. Scrapy is a Python library you embed: you write spiders, define selectors, and run them in your own process. Firecrawl is a service you call: it owns the browser, the queue and the retry logic, and you receive Markdown or JSON over HTTP. With Scrapy you control concurrency and memory directly and pay nothing per page; with Firecrawl you give up that control in exchange for not writing the browser orchestration yourself.

Against a plain headless-browser driver such as Playwright used directly, the difference is scope. Playwright gives you a browser and an API to drive it. Firecrawl gives you a browser plus a queue, a rate limiter, a Postgres-backed job store and an LLM-readable output format. If your pipeline already has a queue and a database, much of what Firecrawl adds is redundant, and you are paying for a layer you already built.

Against hosted scraping APIs generally, the distinguishing property here is the licence and the self-host path. AGPL-3.0 with a docker-compose.yaml in the repository means you can run the whole thing yourself, which most commercial scraping APIs do not offer. That is the trade: you take on the operational burden in exchange for not depending on someone else's uptime or pricing.

Licence, maintenance and what an upgrade actually costs

Firecrawl's core is AGPL-3.0. The practical consequence is that if you modify it and offer it to users over a network, the AGPL's source-disclosure obligation is generally understood to apply to your modified version. Running it internally as a backend service is a different question from shipping it inside a product, and the answer depends on facts a licence file cannot settle. This is not legal advice; if you plan to embed Firecrawl in something you distribute or host for third parties, get the question answered by someone qualified before you build on it. The README also points at a hosted service, which is the usual commercial escape hatch for teams that cannot live with the copyleft terms.

On maintenance, the facts are these: the repository is not archived, and the most recent push recorded is 2026-06-19, which is the same timestamp as the v2.11.0 release. Before that, v2.10 landed on 2026-05-15 and v2.9.0 on 2026-04-10. That is a roughly monthly release cadence through the first half of 2026, followed by a gap. The repository does not explain the gap, and the README does not document a support policy or a release schedule. Draw your own conclusion from the dates rather than from anyone's summary of them.

Upgrade cost depends on which surface you use. SDK users track the firecrawl package and the firecrawl npm package. Self-hosters track the compose file and the image tag. The compose file's default is to build from apps/api rather than pull a pinned image, which means a naive rebuild can move you forward without you choosing to. If you self-host, pin the image and read the release notes for v2.11.0 before rebuilding. The README does not document a rollback procedure or a migration path between versions, so a version bump is a change you should stage rather than apply.

Editorial conclusion

Adopt Firecrawl if you need Markdown or structured JSON out of JS-heavy pages and you are willing to either pay the hosted service or run the API, Redis, Postgres and a Playwright microservice yourself. Do not adopt it if you only need static HTML parsing, since a plain HTTP client plus an HTML parser avoids the whole stack. Verify the AGPL-3.0 obligations against your distribution model, check whether your deployment needs OPENAI_API_KEY set for the Agent and Interact features, and confirm the self-hosted image tag before you pin it.

Frequently asked questions

Why use Firecrawl instead of scraping pages myself?

The README positions it around JavaScript-heavy pages, rotating proxies, rate limits and browser orchestration, plus LLM-ready output in Markdown or structured JSON. If your targets are static HTML, that layer is overhead you do not need.

Is Firecrawl free to use?

The README describes the project as open source and also available as a hosted service, with API keys obtained by signing up at firecrawl.dev. The repository does not state hosted pricing, so the self-hosted Docker Compose route is the only cost structure the repository documents.

Is Firecrawl safe to use?

The README does not address safety or data handling. What the compose file does show is that self-hosting keeps the API, Redis, Postgres and the Playwright microservice inside your own infrastructure, and that USE_DB_AUTHENTICATION defaults to false.

How do I install Firecrawl?

For the hosted service you install the firecrawl SDK or call the REST endpoint with an API key. For self-hosting, the repository ships docker-compose.yaml at the top level and a SELF_HOST.md file, and the compose file builds the API from apps/api by default.

How do I use Firecrawl locally?

The self-host path is the Docker Compose stack: the API plus Redis, a Postgres instance addressed as nuq-postgres, and a Playwright microservice at playwright-service:3000/scrape. Concurrency defaults such as BROWSER_POOL_SIZE of 5 and MAX_CONCURRENT_JOBS of 5 are set through environment variables.

How do I use Firecrawl MCP?

The README gives a JSON block under the mcpServers key with a firecrawl-mcp entry whose command is npx and whose args start with -y firecrawl-mcp. It also documents a one-command skill install, npx -y firecrawl-cli@latest init --all --browser, and states you should restart your agent afterwards.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/firecrawl-firecrawl.svg)](https://hysenlabs.com/projects/firecrawl-firecrawl)
Community notes

Community notes