# htmlparser2: a callback parser that hands the DOM to four sibling packages

> htmlparser2 parses HTML and XML by firing callbacks, fourteen of them, all optional, and it reaches its speed by taking shortcuts the page will name for you. It ships no DOM, no query engine, and no CommonJS entry point: those come from domhandler, css-select, cheerio, and dom-serializer, or from parse5 if you need spec compliance instead of speed.

**fb55/htmlparser2** — The fast & forgiving HTML and XML parser

- Repository: https://github.com/fb55/htmlparser2
- Website: https://feedic.com/htmlparser2
- Stars: 4,790 · Forks: 403
- Language: TypeScript
- License: MIT
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/fb55-htmlparser2

## Speed comes from shortcuts, and the page names the alternative

The project's own framing is blunt about the trade. It calls itself the fastest HTML parser and says in the same sentence that it takes some shortcuts to get there, then points at parse5 for anyone who needs strict HTML spec compliance. That is the whole pitch in two clauses, and it decides which library you want before you write a line. Forgiving parsing means malformed markup still produces events instead of throwing; spec compliance means the tree matches what a browser would build. Choosing htmlparser2 and then discovering your pipeline needs the second behaviour means rewriting the parse step, because the shortcuts are not a flag you can switch off.

## Fourteen callbacks, all optional, and ontext fires more than once

The library hands you a handler object and calls you back. Every callback is optional, so you implement the subset you need: `onopentag` with an aggregated attributes object, `onopentagname` and `onattribute` if you would rather have the tag name before its attributes are parsed, `onclosetag`, `ontext`, `oncomment`, `oncdatastart` and `oncdataend`, `onprocessinginstruction`, `oncommentend`, `onparserinit`, `onreset`, `onend`, and `onerror`. The mechanism is described as consuming documents with minimal allocations. Two details will bite you. `ontext` may fire several times for a single text node, which is why the sample output on the page is labelled as having its text events combined. And `onattribute` reports the quote character as `"`, `'`, `null` for unquoted, or `undefined` when the attribute has no value at all, as in `disabled`.

## xmlMode is one boolean that changes five behaviours at once

The parser ships in HTML mode, and the option table shows what one flag moves. Setting `xmlMode` to `true` treats the document as XML, which affects entity decoding, self-closing tags, CDATA handling, and more; the page says to set it for XML, RSS, Atom, and RDF feeds. Alongside it sit `decodeEntities`, default `true`, which turns `&amp;` into `&`, and `lowerCaseTags` and `lowerCaseAttributeNames`, whose defaults are written as the negation of `xmlMode`. That default is the part worth internalising: in HTML mode you get lowercase tag and attribute names, in XML mode you keep the document's casing, and because the defaults are tied to the flag, explicitly setting one of them will not change the other.

## The parser gives you callbacks; the DOM lives in four other packages

Install it with one command, and try it without writing anything local:

```sh
npm install htmlparser2
```

A live demo runs on AST Explorer. The rest of the ecosystem is a chain of responsibilities, and the page names each link: htmlparser2 parses, domhandler is the handler that turns documents into a DOM, domutils holds utilities for working with that DOM, css-select is the CSS selector engine compatible with it, cheerio is the jQuery API over it, and dom-serializer writes it back out. So the ergonomics you would expect from a parser library, a tree, a querySelector method, a jQuery-shaped wrapper, none of it is in this package. If you only need to walk tags as they stream past, that is exactly right; if you need to query, you are assembling a stack.

## One usage sample, and it stops before the parser is fed anything

The page shows the handler shape and then stops:

```js
import * as htmlparser2 from "htmlparser2";

const parser = new htmlparser2.Parser({
    onopentag(name, attributes) {
        /*
         * This fires when a new tag is opened.
         *
         * If you don't need an aggregated `attributes` object,
         * have a look at the `onopentagname` and `onattribute` events.
         */
        if (name === "script" && attributes.type === "text/javascript") {
            console.log("JS! Hooray!");
        }
    },
    ontext(text) {
```

What follows on the page is the sample output rather than the rest of the code, and that output is the interesting part, because it shows text arriving in fragments:

```
--> Xyz
JS! Hooray!
--> const foo = '<<bar>>';
That's it?!
```

Read the arrows as separate text events. The handler is created here, but the page never shows the call that feeds a document into it, so the surrounding API is something you pick up from the type definitions.

## The manifest is ESM only, with two stream entry points

The package manifest declares `"type": "module"` and every entry in its exports map points at an ES module under dist: the root export, plus `./WebWritableStream` and `./WritableStream` as separate subpaths. There is no CommonJS branch in that map, so a project still on CommonJS cannot require this package and has to reach it through a bundler or a dynamic import. Two other fields matter in practice. `sideEffects` is set to `false`, which is the declaration that lets a bundler drop code it thinks is unused, and the keyword list includes streams, which matches those two dedicated entry points: a consumer who only wants the web stream adapter does not have to pull the main bundle to get it.

## Sources ship beside the build, with the fixtures left out

The published file list is unusual enough to be worth reading: it includes `dist` and `src`, and then excludes `!**/*.spec.ts`, `!src/**/__tests__/**`, `!src/**/__fixtures__/**`, and `!src/**/__snapshots__/**`. So the TypeScript source is available inside node_modules, which is how you check what an edge case actually does instead of guessing from the built output, while the test files, the fixtures, and the snapshots are deliberately kept out of the tarball. That cuts both ways: readable source, but no examples of the malformed input the parser is built to survive. The repository root is small for the same reason, with `src/`, `biome.json`, `eslint.config.mjs`, two tsconfigs, a committed lockfile, and a SECURITY.md.

Badges at the top of the page point at the package on npm, at a coverage report, and at a Node test workflow, so tests and coverage run on the default branch. The release cadence is quick: v10.1.0 on 2026-01-21, v11.0.0 on 2026-03-19, and v12.0.0 on 2026-03-20, two major versions a day apart, with the manifest already at 12.0.0 and the last push dated 2026-10-05.

## Conclusion

Use htmlparser2 when you want to walk a document once, in a stream, without building a tree you do not need, and when forgiving parsing beats spec compliance for your input. Do not reach for it when the input is untrusted markup that has to be validated strictly, since the project itself points to parse5 for that, and remember that any querySelector-shaped API you expect comes from another package. Before you install, check the release you pin, because major versions arrive quickly, and confirm your build can consume an ESM-only package before you add it to a CommonJS codebase.

## FAQ

### What does an HTML parser do, and what does htmlparser2 add?

An HTML parser turns markup into events or a tree. htmlparser2 does the first: it exposes a callback interface with fourteen optional events, described as consuming documents with minimal allocations, and it reaches that speed by taking shortcuts rather than enforcing the HTML spec strictly.

### What is the best HTML parser?

The project's own answer is that it is the fastest and that it takes shortcuts to get there, and it points at parse5 for anyone who needs strict HTML spec compliance. That choice depends on whether forgiving parsing or spec accuracy matters more for your input.

### Does htmlparser2 give me a DOM or a querySelector method?

No, the parser itself delivers callbacks. Turning documents into a DOM is the job of domhandler, querying that DOM is css-select's, working with it is domutils', and the jQuery-shaped API is cheerio's.

### When should I turn on xmlMode in htmlparser2?

Set `xmlMode` to true for XML, RSS, Atom, and RDF feeds. It affects entity decoding, self-closing tags, and CDATA handling, and `lowerCaseTags` and `lowerCaseAttributeNames` default to the opposite of it, so casing follows the mode unless you set them yourself.

### Can a CommonJS project require htmlparser2?

Not through the manifest as published. It declares `"type": "module"` and every export entry points at an ES module under dist, with no CommonJS branch, so you need a bundler or a dynamic import. It also sets `sideEffects` to false and ships `./WebWritableStream` and `./WritableStream` as separate entry points.

## Sources

- [fb55/htmlparser2 on GitHub](https://github.com/fb55/htmlparser2)
- [License: MIT](https://github.com/fb55/htmlparser2/blob/master/LICENSE)
- [Project website](https://feedic.com/htmlparser2)
- [README](https://github.com/fb55/htmlparser2/blob/master/README.md)
- [Releases](https://github.com/fb55/htmlparser2/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/fb55-htmlparser2
