# node-crawler: a TypeScript web spider with Cheerio built in

> The bda-research/node-crawler package wraps got, Cheerio and a priority queue into one crawler object, and v2 dropped CommonJS to become ESM-only. Here is what the README documents, what it leaves out, and who should stay on v1.

**bda-research/node-crawler** — Web Crawler/Spider for NodeJS + server-side jQuery ;-)

- Repository: https://github.com/bda-research/node-crawler
- Stars: 6,794 · Forks: 864
- Language: TypeScript
- License: MIT
- Published: 2026-09-22 · Updated: 2026-09-22 · Language: en
- Canonical page: https://hysenlabs.com/projects/bda-research-node-crawler

## What node-crawler solves for Node.js scraping work

Writing a scraper by hand means assembling four things that have nothing to do with the data you want: an HTTP client, a concurrency limiter, a parser, and a de-duplication layer so you do not fetch the same URL twice. node-crawler is a package that bundles all four behind one constructor. The package.json describes it as a ready-to-use web spider that works with proxies, asynchrony, rate limit, configurable request pools, jQuery, and HTTP/2 support. The README lists the same set as features: server-side DOM with automatic jQuery insertion through Cheerio, a configurable pool size and retries, rate limit control, a priority queue of requests, and charset detection and conversion.

The audience is a Node.js developer who wants to point at a list of URLs and get parsed HTML back inside a callback. It is not a framework. There is no scheduler, no persistent frontier, no dashboard. The unit of work is a Crawler instance with a queue attached, and the queue lives in memory for the life of the process.

## How the queue, the pool and Cheerio fit together

A Crawler instance is constructed with options, and URLs are pushed onto it with add. The README's first example constructs the crawler with maxConnections: 10 and a callback, then calls c.add with a single URL, an array of URLs, or an array of option objects. Each option object can carry its own callback, which suppresses the global one for that request.

The callback signature is (error, res, done), and done() must be called or the queue stalls. That is the single most important mechanical detail in the whole API. Inside the callback, res.$ is a Cheerio instance, described in the README as a lean implementation of core jQuery designed specifically for the server, so selectors like $("title").text() work as they would in a browser.

Two options change the shape of the pool. Setting rateLimit to a number of milliseconds forces maxConnections to 1, so requests are serialised with a minimum gap between them. Setting encoding to null stops the body being converted to a string, which is what you want for images and PDFs. The README also documents a preRequest hook that runs before each request, with the explicit caveat that direct requests do not trigger it. userParams is the documented way to carry state from the request options through to the callback, read back as res.options.userParams.

Underneath, the dependency list shows got for HTTP, hpagent and http2-wrapper for proxy and HTTP/2 transport, iconv-lite for charset conversion, and seenreq for URL de-duplication. The architecture is a thin orchestration layer over those libraries rather than a from-scratch client.

## Installing node-crawler and crawling a page

The README states a hard requirement of Node.js 22 or above, and the package.json engines field agrees. Install with npm:

```bash
npm install crawler
```

The package is ESM-only. The README warns that Crawler v2 was designed as a native ES module and no longer offers a CommonJS export, and recommends converting to ESM. A minimal crawl looks like the README's example:

```js
import Crawler from "crawler";

const c = new Crawler({
    maxConnections: 10,
    callback: (error, res, done) => {
        if (error) {
            console.log(error);
        } else {
            const $ = res.$;
            console.log($("title").text());
        }
        done();
    },
});

c.add("http://www.amazon.com");
```

Run that and you should see the page title printed once the response arrives. If nothing prints, the usual cause is a missing done() call in a custom callback. To slow a crawl down, pass rateLimit: 1000 in the constructor options, which the README notes forces maxConnections to 1 and enforces a minimum 1000 ms gap between tasks.

For binary downloads the README shows encoding: null together with jQuery: false, writing res.body to a stream. The jQuery: false flag there is only to suppress a warning message, not to change the download.

## The ESM-only decision and the migration path back

The largest constraint in v2 is not a feature, it is the module format. The README attributes the change to the dependency migration away from request and toward got, and states plainly that there is no CommonJS export. For a team with a large CommonJS service, this is a blocking change, not an inconvenience.

The README offers one escape hatch: upgrade to v2.0.3-beta with npm install crawler@beta, which supports both ESM and CommonJS builds. That is a beta channel, and the README does not describe a support window or a planned release date for a stable dual-format version. The package.json exports map in the repository lists both a require and an import entry pointing at dist/index.cjs and dist/index.js, which is the build output of the tsup step, so the dual build exists in the toolchain even though the stable release is documented as ESM-only.

There is a second breaking change buried in the same warning: code that used the body parameter to send form data in POST requests must be updated to use form, and the README says this applies even in the beta version. If you have POST-based crawling, that rename is the first thing to grep for.

## Where node-crawler is the wrong tool

The parser is Cheerio, and Cheerio parses HTML strings. It does not execute JavaScript. Any target that renders its content client-side will hand the crawler an empty shell, and no option in the documented API changes that. If your target is a single-page application, a Puppeteer or Playwright based crawler is the right category of tool, and node-crawler is not a substitute.

Memory is the second boundary. The priority queue is in-process, and the README does not document a persistence or resume mechanism. A crawl that dies halfway through a large URL set starts over, minus whatever seenreq has already recorded in that process. The README is also silent on retry counts and backoff configuration beyond listing retries as a feature, so the exact retry semantics are not something you can plan around from the documentation alone.

Finally, the project ships an agent skill under skills/node-crawler/SKILL.md, installable through openclaw skills install node-crawler or by unzipping a release bundle into a skills directory. That is useful if you drive crawling from an AI agent, but it is a wrapper around the same library, not a different capability. It does not make the crawler render JavaScript either.

## Choosing between node-crawler and a Puppeteer-based crawler

The real alternative for most teams is a headless-browser crawler such as one built on Puppeteer. The difference is not speed or ergonomics, it is what the tool can see. A Puppeteer-based crawler boots a Chromium instance, navigates, and waits for the page to finish rendering, so it can read content injected by client-side JavaScript and interact with the page. node-crawler fetches bytes and hands them to Cheerio, so it reads exactly what the server sent.

That trade runs both ways. A browser per worker costs far more memory than an HTTP client, and the concurrency model is entirely different: node-crawler's maxConnections and rateLimit options have no direct equivalent when each unit of concurrency is a browser context. For static HTML, server-rendered pages, sitemaps, and feeds, the Cheerio path is the cheaper and simpler one, and the built-in charset conversion through iconv-lite handles pages that a naive fetch would mangle.

The practical rule: if curl returns the content you want, node-crawler is enough. If curl returns an empty div, no amount of configuration on this package will help.

## Maintenance, licence and upgrade cost

The repository is not archived, and the last push was on 2026-06-18. The most recent tagged release is v2.1.1 from 2026-06-16, preceded by v2.1.0 on 2026-06-04. Before those, the release history jumps back to 0.4.1 in 2014, which tells you the v1 line and the v2 rewrite are separated by a long gap rather than a steady cadence. Treat the v2 line as the one under current work and v1 as legacy.

The licence is MIT, which the repository states in both the LICENSE file and the package.json license field. MIT permits commercial use and modification provided the copyright notice and permission notice are retained. That is the extent of what the repository says; it is not legal advice, and if you redistribute the package inside a product you should confirm the notice requirements with whoever handles licensing on your side.

Upgrade cost is dominated by the ESM migration, not by API churn. The README's Differences and Breaking Changes section is the entry point, and it flags the body-to-form rename for POST requests. The beta channel at crawler@beta is the documented stopgap for teams that cannot move to ESM yet, but the README does not commit to when that channel stops being beta.

## Conclusion

Adopt node-crawler if you are on Node.js 22 or above, already write ESM, and want Cheerio parsing and a request queue in the same package rather than wiring got and cheerio together yourself. Do not adopt it if your codebase is CommonJS and the migration is not scheduled, or if you need the headless-browser rendering that a Puppeteer-based crawler provides. Before committing, check that npm install crawler resolves to a version whose exports map matches your module system, and confirm on the repository that the last push was on 2026-06-18 and that no release has superseded v2.1.1 since.

## FAQ

### Which Node.js version does node-crawler v2 require?

The README states Node.js 22 or above, and the package.json engines field sets node to >=22. Installing on an older runtime is outside what the project documents.

### Does node-crawler work with CommonJS?

No. The README states that Crawler v2 was designed as a native ES module and no longer offers a CommonJS export. It points teams that need both formats at v2.0.3-beta, installed with npm install crawler@beta.

### What is crawling versus scraping in node-crawler terms?

The package covers both halves: the queue and pool handle fetching pages at a controlled rate, while Cheerio, described in the README as a server-side jQuery implementation, handles extracting values from the HTML that comes back.

### Do web crawlers still exist?

Yes. node-crawler is one, published on npm as crawler, and its README documents a v2 line with releases in 2026 alongside a legacy v1 line whose last release was 0.4.1 in 2014.

## Sources

- [bda-research/node-crawler on GitHub](https://github.com/bda-research/node-crawler)
- [Issues](https://github.com/bda-research/node-crawler/issues)
- [License: MIT](https://github.com/bda-research/node-crawler/blob/master/LICENSE)
- [README](https://github.com/bda-research/node-crawler/blob/master/README.md)
- [Releases](https://github.com/bda-research/node-crawler/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/bda-research-node-crawler
