# cheerio: jQuery-style HTML parsing for Node.js scrapers

> cheeriojs/cheerio is an MIT-licensed TypeScript library that gives Node.js and browser code a jQuery-like API over parse5 or htmlparser2. It is a good fit for server-side HTML extraction, and a poor fit when you need a real rendering engine.

**cheeriojs/cheerio** — The fast, flexible, and elegant library for parsing and manipulating HTML and XML.

- Repository: https://github.com/cheeriojs/cheerio
- Website: https://cheerio.js.org
- Stars: 30,497 · Forks: 1,718
- Language: TypeScript
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/cheeriojs-cheerio

## What cheerio is for, and who ends up using it

cheerio solves one narrow problem: taking an HTML or XML string and letting you query and edit it with jQuery syntax, without a browser. The README describes it as implementing "a subset of core jQuery" while removing "all the DOM inconsistencies and browser cruft". That framing is accurate about the trade. You get `.text()`, `.attr()`, `.addClass()`, `.html()` and CSS-style selectors, but you do not get a window, layout, events, or anything that depends on rendering.

The audience is server-side developers. The repository topics include scraper, selector, htmlparser2 and jquery, and the README asks production users to add themselves to a wiki page. Typical work is pulling fields out of a fetched page, rewriting a fragment of markup before it goes into an email or a template, or normalising XML from an API. Anything that needs a headless browser is outside the scope, and the README never claims otherwise.

## How cheerio parses: parse5 by default, htmlparser2 on request

The README states that cheerio wraps around parse5 for HTML parsing and can optionally use the forgiving htmlparser2. That is the central architectural fact. parse5 follows the HTML specification closely, including its error recovery rules, so malformed markup gets fixed in the same way a browser would fix it. htmlparser2 is faster and more tolerant in a different sense: it does not attempt full spec-compliant tree construction.

The package layout reflects this. package.json declares `"type": "module"` and an exports map with separate `import` and `require` conditions, plus browser variants and a `./slim` entry point. The main entry is the full build; `./slim` is the lighter one that skips the default parser wiring. If you are bundling for the browser, the exports map resolves to `./dist/esm/browser/index.js`, so the same source can target both environments.

Loading is explicit. In jQuery the document is implicit because jQuery runs inside a page. In cheerio you pass the HTML in, which is why every example starts with `cheerio.load(...)`. Selections then work against that loaded document, and `$.root().html()` renders the whole thing back out. The README also documents a DOM-node-shaped object with `tagName`, `parentNode`, `previousSibling`, `nextSibling`, `nodeValue`, `firstChild`, `childNodes` and `lastChild`, so code that walks the tree manually has stable property names to rely on.

## Installing cheerio and running a first extraction

The README lists npm, yarn, bun and Deno as supported package managers and gives the commands directly. For npm:

```bash
npm install cheerio
```

Bun and Deno users get their own lines in the same block:

```bash
bun add cheerio
# or
deno install cheerio
```

Once installed, the README's opening example is the smallest complete program. It loads a fragment, edits it, and renders it back:

```js
import * as cheerio from 'cheerio';
const $ = cheerio.load('<h2 class="title">Hello world</h2>');

$('h2.title').text('Hello there!');
$('h2').addClass('welcome');

$.html();
//=> <html><head></head><body><h2 class="title welcome">Hello there!</h2></body></html>
```

Two things are worth noticing in that output. First, cheerio normalised the fragment into a full document with `html`, `head` and `body` elements. Second, `addClass` merged into the existing class attribute rather than replacing it. For a read-only extraction, the selector and accessor forms are what you want. The README shows scoped selectors with a context argument, and attribute and text accessors:

```js
$('.apple', '#fruits').text();
//=> Apple

$('ul .pear').attr('class');
//=> pear

$('li[class=orange]').html();
//=> Orange
```

The second argument to `$` narrows the search to a context, which matters when a page repeats the same class names in several regions. To get the markup of a single matched element rather than the whole document, the README points at the `outerHTML` prop: `$('.pear').prop('outerHTML')` returns `<li class="pear">Pear</li>`. For plain text, `$('body').text()` on `This is <em>content</em>.` returns `This is content.`, with tags stripped.

## Where cheerio stops: no rendering, and parser choice is not free

The most common wrong turn is expecting cheerio to handle a page that builds its content in the browser. cheerio parses a string. If the HTML you fetched contains an empty container and a script tag that fills it, cheerio will faithfully return the empty container. There is no JavaScript execution, no network waterfall, no layout. That is a hard boundary, and the README does not present it as anything else.

The second limitation is subtler and comes from the parser split. Because parse5 and htmlparser2 recover from broken markup differently, the same input can produce different trees depending on which one is active. The README describes htmlparser2 as "forgiving", which is exactly the property that makes its output diverge from spec-compliant parsing on damaged documents. If your selectors were written against one parser and the build switches to the other, results can change without any code change on your side. The exports map makes this easy to do accidentally, since the main entry and `./slim` resolve to different files.

There is also a cost to the jQuery surface itself. cheerio implements a subset, not all of it, so a selector or method you rely on from jQuery may simply not exist. The README says "subset" and leaves the enumeration to the API documentation on cheerio.js.org rather than listing exclusions inline.

## cheerio versus a headless browser

The realistic alternative for scraping work is a headless browser driven by something like Playwright or Puppeteer. The difference is not speed or API taste, it is where the work happens. A headless browser loads the page, executes scripts, applies styles, and then hands you a live DOM. cheerio takes a string you already have and builds a static tree from it.

That means the browser route handles client-rendered pages, login flows, and anything gated behind script execution, at the cost of a much heavier process and a real browser binary in your deployment. The cheerio route handles server-rendered HTML and XML with a fraction of that footprint, and the README's screencast section explicitly frames cheerio as a replacement for JSDOM plus jQuery. If your target returns complete markup in the first response, a browser is overhead you do not need. If it does not, cheerio cannot help you, regardless of how the selectors read.

## Licence, maintenance and what an upgrade costs

cheerio is MIT licensed, stated in both the README badge area and the `license` field in package.json. MIT is permissive: you can use it commercially, modify it and redistribute it, provided the copyright notice and permission notice travel with it. That is the general shape of the licence, not legal advice for your situation; if you vendor or patch the source, read LICENSE in the repository root.

The repository is not archived, and the last push was on 2026-09-10. The most recent release listed is v1.2.0 from 2026-01-23, preceded by v1.1.2 on 2025-07-21 and v1.1.1 on 2025-07-20. The gap between the 1.1.x pair and 1.2.0 is roughly six months, so minor releases are not frequent, and the 1.1.1 to 1.1.2 turnaround of one day suggests patch releases come out when something needs fixing rather than on a schedule.

For upgrade planning, the practical risk is the parser question rather than the version number. The package ships dual ESM and CommonJS builds through the exports map, so a major Node.js or bundler change can affect which file you actually load. The repository carries a benchmark directory and a vitest config, which means the project maintains its own performance and test harness; running your own selectors against both the main entry and `./slim` is the cheapest way to find out whether a bump changes your output.

## Conclusion

Adopt cheerio if you are extracting data from server-rendered HTML or XML in Node.js and want jQuery selectors without a browser. Do not adopt it if the page requires JavaScript execution before the content appears, or if you need CSS layout and computed styles; cheerio parses and manipulates a document, it does not render one. Before committing, verify which parser your build pulls in by checking the exports map in package.json and confirming whether you are importing the main entry or ./slim, and confirm that the selectors you rely on behave the same on malformed markup with parse5 as they do with htmlparser2.

## FAQ

### How to install cheerio?

Install it with a package manager such as npm, yarn, bun or Deno. The README gives `npm install cheerio`, `bun add cheerio` and `deno install cheerio` as the supported forms.

### How to use cheerio in Node.js?

Import the module, pass an HTML string to `cheerio.load`, then use jQuery-style selectors on the returned function. The README's first example loads a heading, changes its text with `.text()` and its class with `.addClass()`, then renders the document with `$.html()`.

### How to use cheerio?

The workflow is load, select, then read or modify. The README shows scoped selectors such as `$('.apple', '#fruits').text()` and attribute reads such as `$('ul .pear').attr('class')`, with `$.root().html()` used to render the result back to a string.

### How to use cheerio js?

cheerio exposes a subset of core jQuery over a document you supply, so the same selector syntax you know from jQuery applies. It parses with parse5 by default and can optionally use htmlparser2, and it runs in both Node.js and the browser according to the README.

## Sources

- [cheeriojs/cheerio on GitHub](https://github.com/cheeriojs/cheerio)
- [License: MIT](https://github.com/cheeriojs/cheerio/blob/main/LICENSE)
- [Project website](https://cheerio.js.org)
- [README](https://github.com/cheeriojs/cheerio/blob/main/README.md)
- [Releases](https://github.com/cheeriojs/cheerio/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/cheeriojs-cheerio
