Crawlee: a Node.js scraping library that swaps HTTP, Cheerio, Playwright and Puppeteer behind one interface
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
At a glance
- What is it?
- Crawlee is an Apache-2.0 TypeScript library for building crawlers in Node.js, with a persistent request queue, pluggable storage and proxy rotation. The code is pushed frequently, but the README does not document rollback and the v4 line is still at a release candidate.
- Who is it for?
- Adopt Crawlee if your team writes JavaScript or TypeScript and wants one crawler abstraction that can start on raw HTTP and move to a real browser without rewriting the request handler. Do not adopt it if you need a Python codebase, since that is a separate repository, or if you want a finished extraction product rather than a library you assemble.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 12 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
What Crawlee actually solves for Node.js scraping teams
Writing a scraper by hand means writing the same supporting code every time: a queue that survives a restart, retry logic, concurrency limits, a place to put extracted rows, and a way to rotate proxies when a site starts returning 403. Crawlee packages that infrastructure as a library. The README calls it "a web scraping and browser automation library" and says it "covers your crawling and scraping end-to-end".
The intended audience is developers working in JavaScript or TypeScript who already know they need to crawl something. The repository is a pnpm workspace with a packages/ directory and a Lerna config, so the published `crawlee` package is assembled from scoped packages such as `@crawlee/core`, `@crawlee/types` and `@crawlee/utils`. For a Python team, this repository is the wrong one: the README points to a separate project, Crawlee for Python, at github.com/apify/crawlee-python.
The more interesting claim is the abstraction. Crawlee's feature list promises a "single interface for HTTP and headless browser crawling", with Playwright and Puppeteer usable "with the same interface". In practice that means the shape of your request handler stays roughly constant while the engine underneath changes, which is the part that saves real work when a target site turns out to render everything client-side.
How the crawler, queue and storage fit together
The core object is a crawler instance. You construct it with a requestHandler callback, then call `crawler.run([...urls])` with the seed URLs. The crawler pulls a URL from its queue, executes the handler with a context object, and repeats until the queue drains. The README example destructures `request`, `page`, `enqueueLinks` and `log` from that context, which tells you what the handler is expected to do: read the current page, extract data, and optionally push more URLs back into the queue.
`enqueueLinks()` is the mechanism that turns a single-page scraper into a crawl. It extracts links from the current page and adds them to the queue, so breadth-first and depth-first traversal come from the queue rather than from your own recursion. The README lists the queue as persistent and says it supports both breadth and depth first ordering.
Storage is split in two. Tabular results go through `Dataset.pushData()`, which the README says writes JSON to `./storage/datasets/default` by default. Files and other artifacts are handled separately by the result storage layer. Because both are pluggable, the same handler code can write to local disk during development and to a cloud target in production. The default location is `./storage` in the current working directory, overridable through Crawlee configuration.
Around that core sit the pieces that keep long crawls alive: proxy rotation and session management, configurable routing, error handling and retries, and hooks for customising lifecycles. Scaling is described as automatic based on available system resources.
Installing Crawlee and running a first crawler
Crawlee requires Node.js 16 or higher according to the README. The fastest path is the CLI, which the README describes as installing dependencies and adding boilerplate. It scaffolds a project and then you start it:
npx crawlee create my-crawler
cd my-crawler
npm startIf you would rather add Crawlee to an existing project, install it alongside a browser automation library. Playwright is not bundled, and the README gives the reason plainly: it keeps the install size down.
npm install crawlee playwrightThe README's worked example builds a `PlaywrightCrawler` that logs each page title, pushes a `{ title, url }` record into the dataset, and calls `enqueueLinks()` to continue the crawl. The `headless: false` option is present but commented out, which is how you watch the browser window during debugging.
import { PlaywrightCrawler, Dataset } from 'crawlee';
const crawler = new PlaywrightCrawler({
async requestHandler({ request, page, enqueueLinks, log }) {
const title = await page.title();
log.info(`Title of ${request.loadedUrl} is '${title}'`);
await Dataset.pushData({ title, url: request.loadedUrl });
await enqueueLinks();
},
});
await crawler.run(['https://crawlee.dev']);After the run finishes, look in `./storage/datasets/default` for the JSON records. If you want to test unreleased fixes, the README documents installing the automated beta builds with `npm install crawlee@next`. If you also use the Apify SDK, the README warns that you need dependency overrides in `package.json` for `@crawlee/core`, `@crawlee/types` and `@crawlee/utils` so you do not end up with two Crawlee versions installed.
Where Crawlee stops being the right tool
The abstraction has a cost. Because Crawlee is a library rather than a finished scraper, you still own the extraction logic for every site, the selector maintenance when markup changes, and the decision about which engine to use. Nothing in the README suggests Crawlee detects that a page needs JavaScript rendering and switches engines for you; choosing between HTTP, Cheerio, JSDOM, Playwright and Puppeteer is a decision you make when you construct the crawler.
Browser crawling is also heavier than it looks. The README treats Playwright as a separate install precisely because bundling it would bloat the package. Running a headless browser per worker consumes far more memory than raw HTTP requests, so the "automatic scaling with available system resources" behaviour will settle at a much lower concurrency for browser crawlers than for HTTP ones. If your target is a JSON API, reaching for a browser engine is wasted capacity.
The bot-evasion claims deserve care. The README states that crawlers "will appear human-like and fly under the radar of modern bot protections even with the default configuration", and lists browser-like headers, TLS fingerprint replication and human-like fingerprints. That is a description of defaults, not a guarantee against any specific anti-bot system, and the README does not name which protections have been tested. Treat it as a starting point that reduces obvious fingerprinting, not as a promise of access.
Finally, the version situation is genuinely awkward for production. The most recent release is v4.0.0-rc.0, a release candidate, while v3.18.1 is the stable line. A repository that ships a release candidate alongside a stable minor means you must decide deliberately which one you pin. The README does not document rollback behaviour for a bad upgrade, and MIGRATIONS.md exists in the repository precisely because majors move things.
How Crawlee differs from Scrapy and from raw Playwright
The most common comparison is Crawlee versus Scrapy, and the split is mostly about language and runtime. Scrapy is a Python framework with its own engine, scheduler and item pipeline, and it expects you to work inside its structure. Crawlee is a TypeScript library that you import into an ordinary Node.js program, so your existing tooling, your package manager and your deployment target stay in place. If your team writes Python, the relevant comparison is not this repository at all but Crawlee for Python, which the README links separately. The two Crawlee implementations are distinct projects, so a JavaScript answer does not transfer to a Python codebase.
The second comparison, Crawlee versus Playwright on its own, is about what surrounds the browser. Playwright gives you a browser and a page API. Crawlee gives you a queue, a dataset, retries, proxy rotation and sessions wrapped around that page API, and lets you swap Playwright for Puppeteer, Cheerio or plain HTTP without changing the handler contract. If your job is a single page rendered by JavaScript and you already have somewhere to put the output, Playwright alone is less machinery. If your job is thousands of URLs with restarts and proxy changes, the queue and storage layers are the reason to add Crawlee.
Both comparisons point at the same design choice: Crawlee is deliberately unopinionated about extraction and opinionated about orchestration.
Maintenance, licensing and the cost of upgrading
The repository is not archived, and the last push was on 2026-09-10, so development is current. The release cadence visible in the notes is steady: v3.18.0 on 2026-08-04, v3.18.1 on 2026-08-12, and v4.0.0-rc.0 on 2026-08-13. The README also states that automated beta builds are produced for every merged code change and published to npm, so the `next` tag moves far more often than the stable tag does.
That cadence is the upgrade cost. Running `crawlee@next` means tracking a moving target, and the README's own advice about dependency overrides for the Apify SDK shows how easily a project ends up with two copies of the same scoped package. Pin an exact version, and read MIGRATIONS.md before crossing a major boundary. The repository also ships RELEASE.md and a CHANGELOG.md, which are the places to check what a version bump actually changes.
Licensing is Apache-2.0, declared both in the repository's package.json and in LICENSE.md. Apache-2.0 is a permissive licence that includes an explicit patent grant, which matters if you are embedding the library in a commercial product. This is a description of the licence text, not legal advice; if your organisation has a policy on permissive licences or on the patent clause, route it through whoever normally reviews that. Note that the licence covers Crawlee itself, not Playwright, Puppeteer or any other dependency you install alongside it, each of which carries its own terms.
Editorial conclusion
Adopt Crawlee if your team writes JavaScript or TypeScript and wants one crawler abstraction that can start on raw HTTP and move to a real browser without rewriting the request handler. Do not adopt it if you need a Python codebase, since that is a separate repository, or if you want a finished extraction product rather than a library you assemble. Before committing, verify which Crawlee major you are pinning, because v4.0.0-rc.0 is a release candidate while v3.18.1 is the stable line, and read MIGRATIONS.md in the repository before upgrading.
Frequently asked questions
What is Crawlee?
Crawlee is a web scraping and browser automation library for Node.js, published as the `crawlee` npm package and written in TypeScript. It provides a single interface for HTTP and headless browser crawling, with a persistent request queue and pluggable storage.
What does a crawler do?
In Crawlee, a crawler pulls URLs from a persistent queue, runs your requestHandler against each page, and uses `enqueueLinks()` to extract links from the current page and add them back to the queue. It repeats until the queue drains.
Is Crawlee free to use?
Yes. The repository declares the Apache-2.0 licence in both package.json and LICENSE.md, which permits commercial use and includes an explicit patent grant. Dependencies you install alongside it, such as Playwright, have their own licences.
What is the difference between crawling and scraping?
The README separates the two jobs: Crawlee's queue and `enqueueLinks()` handle crawling, meaning following links from page to page, while `Dataset.pushData()` handles scraping output by storing extracted records as JSON in `./storage/datasets/default` by default.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/apify-crawlee)
Community notes