# AnyCrawl: a self-hosted Node.js crawler that returns LLM-ready Markdown and SERP data

> AnyCrawl is a TypeScript crawler and scraping toolkit that turns pages into Markdown or JSON and pulls structured search results, with a Docker Compose deployment and a REST API. Its value is the API surface and the engine split; the cost is a Redis, database and browser stack you now operate.

**any4ai/AnyCrawl** — AnyCrawl 🚀: A Node.js/TypeScript crawler that turns websites into LLM-ready data and extracts structured SERP results from Google/Bing/Baidu/etc. Native multi-threading for bulk processing.

- Repository: https://github.com/any4ai/AnyCrawl
- Website: https://anycrawl.dev
- Stars: 3,464 · Forks: 368
- Language: TypeScript
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/any4ai-anycrawl

## The gap AnyCrawl targets: raw HTML is not what a model reads

A crawler that returns raw HTML solves the fetching problem and creates a parsing problem. Anyone building a retrieval pipeline over documentation, product pages or news archives has to strip navigation, scripts and boilerplate before the text is worth embedding. AnyCrawl takes the position that conversion belongs inside the crawler. The README describes it as a toolkit that turns websites into LLM-ready data, and the repository topics list html-to-markdown and rag alongside scrape and serp.

The second half of the pitch is search results. AnyCrawl exposes SERP crawling across Google, Bing, Baidu and other engines through the same API, so a pipeline that needs both page bodies and ranked result lists does not have to integrate two clients. The audience is TypeScript and Node.js teams that want an HTTP service rather than a library import, and that are comfortable running Redis and a browser engine next to their application. If you only ever fetch one page at a time and parse it yourself, this is more infrastructure than the task needs.

## How AnyCrawl is put together: API, Redis queue, scrape workers, browser pool

The repository is a pnpm workspace driven by Turborepo. Top-level entries include apps/, packages/ and turbo.json, and the root package.json defines scripts such as build:api, start:api, start:worker and test:browser-score. The Dockerfile copies package manifests for apps/api and for the packages libs, scrape, search, ai, db and template-client, which tells you the split: an API service, a scraping worker package, a search package and an AI extraction package.

At runtime the pieces talk through Redis. The docker-compose.yml defines an api service on port 8080 and separate scrape workers, including a scrape-puppeteer service, all sharing a network and waiting on redis. The common environment block sets ANYCRAWL_REDIS_URL to redis://redis:6379 and points the API at a SQLite file at /usr/src/app/db/database.db. Work arrives at the API, is queued, and is picked up by workers that drive a browser pool. That pool is governed by explicit keys: ANYCRAWL_BROWSER_MAX_PAGES_PER_BROWSER defaults to 500, ANYCRAWL_BROWSER_MAX_OPEN_PAGES_PER_BROWSER defaults to 20, ANYCRAWL_BROWSER_IDLE_RETIRE_SECS defaults to 3600, and ANYCRAWL_BROWSER_ISOLATE_CONTEXTS defaults to true. Those four values are the real tuning surface for memory and throughput, more than the concurrency numbers.

The engine split matters for cost. The scrape endpoint accepts cheerio, playwright or puppeteer. Cheerio parses static HTML and the parameter table calls it the fastest option; the other two render JavaScript. Choosing cheerio where it works avoids launching a browser entirely, which is the difference between a cheap request and an expensive one.

## Installing AnyCrawl with Docker Compose and making a first scrape call

The README points to docs.anycrawl.dev for full documentation and ships a docker-compose.yml at the repository root for a self-hosted run. Authentication is off by default: the compose environment sets ANYCRAWL_API_AUTH_ENABLED to false, and .env.example shows the same default alongside ANYCRAWL_API_CREDITS_ENABLED=false. If you turn authentication on, generate a key inside the running container. The README gives both the compose and the single-container form:

```bash
docker compose exec api pnpm --filter api key:generate
docker compose exec api pnpm --filter api key:generate -- default
```

The command prints a uuid, a key and credits. Use the printed key as a Bearer token. With the stack up, the API listens on port 8080, which the compose file maps as "8080:8080". A first scrape is a single POST. The README gives this example, with the public endpoint shown and the note that self-hosters should replace it with their own server URL:

```bash
curl -X POST http://localhost:8080/v1/scrape \
  -H 'Content-Type: application/json' \
  -H 'Authorization: Bearer YOUR_ANYCRAWL_API_KEY' \
  -d '{
  "url": "https://example.com",
  "engine": "cheerio"
}'
```

Expect the response to carry the extracted content for that URL. Two request fields are worth setting from the start. max_age controls cache reads in milliseconds, where 0 forces a refresh and a positive value accepts cached content within that age. store_in_cache defaults to true, so repeated calls against the same URL will not refetch unless you set max_age to 0. The README also points to a hosted Playground for generating request code in your preferred language, and to docs/cache.md for self-hosted cache details.

## Proxies, stealth mode and the keys that decide whether a crawl finishes

Anti-bot handling is where self-hosted crawlers usually stall, and AnyCrawl's configuration shows the seams. .env.example exposes ANYCRAWL_PROXY_URL and a separate ANYCRAWL_PROXY_STEALTH_URL, with ANYCRAWL_PROXY_STEALTH_CREDITS defaulting to 5. Timeouts are split by mode: ANYCRAWL_STEALTH_TIMEOUT_MS defaults to 120000 and ANYCRAWL_BASE_TIMEOUT_MS to 60000. The stealth path is slower by design, and the comment in the file states that the Cloudflare Turnstile solver runs in stealth proxy mode only, configured through ANYCRAWL_2CAPTCHA_API_KEY against https://api.2captcha.com.

Sticky proxy sessions are off unless you enable them. The comment in .env.example is unusually specific: all browser proxies must use a {sessionId} placeholder and support at least the declared window from first use, and enabling the feature does not change the provider's TTL. There are no per-proxy overrides. In practice that means a proxy provider whose session format does not match the placeholder will not work with sticky mode, and you should test that before building a pipeline around it.

GeoIP adds another failure point. ANYCRAWL_BROWSER_GEOIP defaults to true, and the comment states that a failed GeoIP resolution fails the browser launch explicitly. A proxy that hides its exit region will not degrade gracefully here; the launch fails. That is a defensible choice because it prevents silently scraping from the wrong country, but it is a hard stop you need to plan for. The concurrency keys carry a warning of their own: .env.example marks ANYCRAWL_MAX_CONCURRENCY and ANYCRAWL_MIN_CONCURRENCY, both defaulting to 50, with the note not to configure them unless you know what you are doing.

## Where AnyCrawl is the wrong tool

The README documents how to start the stack and how to call the API. It does not document what happens to an in-flight crawl when a worker restarts, and no resume or checkpoint behaviour is described anywhere in the documentation. For a one-off extraction that is irrelevant. For a long traversal of a large site, an interrupted run is an unknown quantity, and you should treat resumability as unverified until you find it in the docs or read the worker code.

The deployment shape is the second constraint. A working instance needs Redis, a database and at least one browser-capable worker. On a small machine, a single-process Python crawler that runs in the same interpreter as your script is simpler to operate and simpler to debug. AnyCrawl is the better fit when several services need to share one crawling endpoint, when you want an HTTP boundary between your application and the fetching layer, or when the SERP extraction is part of the same job as the page extraction.

There is also a caching trap. Because store_in_cache defaults to true, a pipeline that calls the same URLs on a schedule will quietly serve stale content unless it sets max_age deliberately. That is a sensible default for cost and a poor default for freshness, and the parameter table is the only place it is explained in the README.

## How AnyCrawl differs from Crawl4AI and Crawlee

Crawl4AI is the closest comparison in purpose, and the difference is the interface. Crawl4AI is a Python library you import and drive from your own process, which means no Redis, no separate API service and no container orchestration; the crawl lives and dies with your script. AnyCrawl inverts that. It runs as a service with a REST API, a queue and workers, so several callers can submit jobs to one deployment and the browser pool is shared and tuned centrally. If your stack is Python and the crawl is part of a batch job, Crawl4AI's model removes an entire layer of infrastructure. If your stack is TypeScript and the crawl is a shared capability, AnyCrawl's model is the one that scales past a single script. The trade is operational: AnyCrawl asks you to run and monitor more, and the README does not document rollback or resume.

Crawlee is closer to a different layer. It is the Node.js crawling and browser automation library, and the AnyCrawl repository depends on it: the compose environment sets ANYCRAWL_CRAWLEE_STORAGE_DIR for the puppeteer worker. Crawlee gives you request queues, storage and browser management inside your own application. AnyCrawl wraps that machinery in an API, adds Markdown conversion for LLM consumption, adds SERP extraction across engines, and adds an AI extraction package for structured JSON output. Choosing Crawlee means writing and owning the service; choosing AnyCrawl means accepting its API shape and its configuration keys in exchange for not writing the service.

## Licence, upgrade cost and what the version numbers imply

AnyCrawl is MIT licensed, and the Dockerfile carries the matching label org.opencontainers.image.licenses=MIT. MIT permits commercial use and modification, and it places no obligation on how you distribute your own application. The practical implication is not legal but operational: nothing stops you from forking the worker to add checkpointing, and if the upstream project changes direction there is no copyleft or relicensing mechanism to stop you from continuing on your own branch. That is a genuine hedge for a component you plan to put in a data pipeline.

The release history is short and recent. v1.0.0 was tagged on 2026-09-09, two days after v1.0.0-beta.37 on 2026-09-07, with v1.0.0-beta.36 before that on 2026-08-11. The last push to the repository was on 2026-09-09. A 1.0 tag that lands two days after a beta is a signal about process, not stability: the API surface is young, and the configuration keys in .env.example are numerous enough that a minor upgrade can plausibly change defaults. Pin the image tag rather than tracking a moving branch, and diff .env.example against your own environment file on every upgrade, because that file is where new keys appear. The repository also ships docker-compose.pg.yml alongside docker-compose.yml, so if SQLite is not what you want in production, a Postgres variant exists at the root without further documentation in the README.

## Conclusion

Adopt AnyCrawl if you want a self-hosted HTTP endpoint that returns Markdown or structured JSON and you are willing to run Redis, a database and a browser engine alongside it. Do not adopt it if you need a crawl that survives a restart, because the README does not document resume behaviour, or if a single-process Python library already covers your job. Before committing, verify the engine you need is listed in ANYCRAWL_AVAILABLE_ENGINES, check that your proxy provider supports the sticky placeholder format, and confirm the cache location under ANYCRAWL_LOCAL_STORAGE_DIR suits your retention rules.

## FAQ

### What is the best tool for web crawling?

There is no single answer, and AnyCrawl's own design shows why: it offers cheerio for static HTML, which the parameter table calls the fastest option, and playwright or puppeteer for pages that need JavaScript rendering. The right choice depends on whether your targets render client-side and whether you want a self-hosted API or a library.

### Do web crawlers still exist?

Yes. AnyCrawl is one, and the README lists site crawling for full-site traversal and collection as a distinct capability alongside single-page scraping. The repository also exposes SERP crawling across multiple search engines.

### Is ChatGPT a web crawler?

ChatGPT is not a crawler. AnyCrawl is the kind of tool that sits on the other side of that line: it fetches pages, converts them to LLM-ready Markdown or structured JSON, and the README describes an AI extraction package for pulling structured data from pages.

### What is the difference between a crawler and a scraper?

In AnyCrawl's own API the distinction is visible: the scrape endpoint takes a single url and returns that page's content, while the README lists site crawling as a separate capability for full-site traversal and collection. One addresses a page, the other addresses a set of pages.

## Sources

- [any4ai/AnyCrawl on GitHub](https://github.com/any4ai/AnyCrawl)
- [License: MIT](https://github.com/any4ai/AnyCrawl/blob/main/LICENSE)
- [Project website](https://anycrawl.dev)
- [README](https://github.com/any4ai/AnyCrawl/blob/main/README.md)
- [Releases](https://github.com/any4ai/AnyCrawl/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/any4ai-anycrawl
