Library / SDK
servo/html5ever avatar
servo/html5ever

html5ever: a Rust HTML5 parser built for browser-grade parsing

High-performance browser-grade HTML5 parser

2,633 stars289 forksRustNOASSERTION

At a glance

What is it?
html5ever is the HTML parser from the Servo project, written in Rust and aimed at WHATWG-conformant parsing plus serialization. It is a library for people building browsers, scrapers or tooling, not a drop-in DOM.
Who is it for?
Adopt html5ever if you are writing Rust and need spec-shaped HTML parsing, either inside a browser engine or in a tool where you control the tree representation. Do not adopt it expecting a DOM, an XPath layer or a Python binding; the README states bindings for Python and other languages are desired, not delivered.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 15 days ago.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 24, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem html5ever solves: HTML that browsers actually accept

HTML in the wild is not well formed. Browsers have a specified recovery algorithm for unclosed tags, mis-nested elements and stray text, and that algorithm is written down in the WHATWG HTML standard. A parser that ignores it will produce a tree that differs from what a browser builds, which matters when you are rendering pages, diffing DOMs or reproducing a layout bug. html5ever is the Servo project's implementation of that algorithm in Rust. The README describes it as parsing and serializing HTML according to the WHATWG specs, with the caveat that some behaviour still differs and most of those differences are tracked in the bug tracker under the web-compat label. The intended audience is narrow and technical: people building a browser engine, or a tool that needs browser-shaped parsing and is written in Rust. If you only need to pull text out of a page, this is more machinery than you need.

Callbacks instead of a DOM: how html5ever is structured

The design decision that shapes everything else is stated plainly in the README: html5ever uses callbacks to manipulate the DOM, therefore it does not provide any DOM tree representation. The parser drives a tokenizer, and as tokens are produced it calls into a TreeSink implementation that you write. You decide what nodes are, how they are allocated and how they are addressed. That is why the repository ships rcdom, a reference DOM built on reference-counted nodes, as a separate crate rather than as part of the parser. The workspace layout reflects the split: html5ever holds the parser, markup5ever holds shared markup infrastructure, web_atoms holds the atom tables, tendril holds the string type, and xml5ever is a sibling parser for XML. Strings are UTF-8 only, and the README notes that other document encodings are a future item handled by converting input. The README also states the goal of passing all html5lib tests while providing the hooks a production browser needs, and names document.write as an example of such a hook.

Installing html5ever and running a first parse

The README's only install instruction is to add the crate as a dependency. The command it gives is cargo add html5ever. From there the README points at examples/html2html.rs, examples/print-rcdom.rs and the API documentation on docs.rs rather than walking through a parse. In the repository those examples live under rcdom/examples/, and rcdom is a workspace member, so the example pulls in the reference DOM in addition to the parser.

bash
cargo add html5ever

To work on the parser itself rather than consume it, the README gives two more commands. The test suite is a git submodule, so it has to be initialised before the tests can run, and local documentation is built with cargo doc into target/doc/.

bash
git submodule update --init
cargo doc

What you should expect after cargo doc is a local copy of the API documentation under target/doc/ in the repository root. The README does not document a standalone CLI, so there is no binary to run against an HTML file; the first real use is writing a TreeSink and calling the parser from your own Rust code, following the two examples named above.

Where html5ever is the wrong tool

The clearest limitation is the one the README states outright: no DOM tree representation. If you want to load a document, query it and walk away, you are writing the tree yourself or depending on rcdom, which the README presents as an example crate rather than the point of the project. The second limitation is conformance. The README says html5ever passes all tokenizer tests from html5lib-tests, with most tree builder tests passing outside the unimplemented features, and that the goal is to pass all of them. That is an honest description of a parser that is close to the spec but not finished against it, and the README directs readers to the web-compat issues for the differences in actual behaviour. Third, the syntax is not XML. The README warns that for correct parsing of XHTML you should use an XML parser, noting that many XHTML documents in the wild are serialized in an HTML-compatible form. If your input is genuinely XHTML, xml5ever in the same workspace is the sibling to look at. Finally, the README lists bindings for Python and other languages as much desired, which means they are not part of this repository. If you are not writing Rust, this is not the parser you install.

How html5ever differs from a general-purpose HTML parser

Most HTML parsing libraries in dynamic languages hand you a tree as the product. You call a function with a string and get back nodes you can query. html5ever inverts that: the tree is your input to the parser, in the form of a TreeSink implementation, and the parser's job is to drive that sink according to the spec's tree construction rules. The practical difference shows up in what you get for free. A tree-first library gives you a query API and often CSS or XPath selection on top. html5ever gives you the parsing algorithm and nothing above it, which is why the README points at examples rather than a tour. This is the same trade-off Servo makes elsewhere: the parser has to serve a browser, where the tree is the browser's own DOM and the parser cannot impose one. For a scraper, a tree-first library is faster to start with. For an engine or a tool that already has a node representation, a callback parser avoids a conversion step between two trees, and that is the case html5ever is built for.

Maintenance, release cadence and the licence files

The repository is not archived, and the last push was on 2026-09-14. The most recent release listed is html5ever-v0.35.0 from 2025-07-02, while the workspace manifest declares version 0.40.1, so the crates in the tree are ahead of the newest tagged release named in the repository's release list. The workspace pins rust-version to 1.85 and edition 2021, which sets a floor for your toolchain. The README notes that html5ever builds against stable Rust, though some optimizations are only supported on nightly releases, so a stable build is expected to work but may not be the fastest configuration. Upgrade cost is mostly the usual Rust dependency work, with one wrinkle: html5ever, markup5ever, web_atoms, xml5ever and tendril are versioned together in the workspace and referenced by path, so a consumer tracking the published crates should expect several of them to move at once. On licensing, the workspace manifest states MIT OR Apache-2.0 and the repository root carries LICENSE-MIT and LICENSE-APACHE, while the repository metadata reports the licence as NOASSERTION. Those two do not agree, so read the LICENSE files in the repository before you rely on either. This is a description of what the files say, not legal advice.

Editorial conclusion

Adopt html5ever if you are writing Rust and need spec-shaped HTML parsing, either inside a browser engine or in a tool where you control the tree representation. Do not adopt it expecting a DOM, an XPath layer or a Python binding; the README states bindings for Python and other languages are desired, not delivered. Before committing, check the web-compat issues in the bug tracker, confirm your Rust toolchain meets the workspace's rust-version of 1.85, and read rcdom/examples/html2html.rs to see how much tree code you will have to supply yourself.

Frequently asked questions

What is html5ever used for?

It parses and serializes HTML according to the WHATWG specs, and the README describes it as an HTML parser developed as part of the Servo project. It is aimed at uses that need browser-shaped parsing, such as a browser engine itself, rather than at simple text extraction. The README states the goal is to pass all html5lib tests while providing the hooks a production browser needs, naming document.write as an example.

How do I install html5ever in a Rust project?

The README gives one command, cargo add html5ever, which adds it as a dependency. It then points at examples/html2html.rs, examples/print-rcdom.rs and the API documentation rather than providing a longer setup guide. The examples live under rcdom/examples/ in the repository.

Does html5ever provide a DOM tree?

No. The README states that html5ever uses callbacks to manipulate the DOM and therefore does not provide any DOM tree representation. The rcdom crate in the same workspace supplies a reference DOM if you want one to build on.

Can I use html5ever from Python?

Not from this repository. The README says bindings for Python and other languages are much desired, which means they are not shipped here. The parser is written in Rust and the documented entry point is the Rust crate.

What is the best HTML parser?

There is no answer to this in the README, and the choice depends on your language and whether you need a tree or a callback interface. What the README does say is that html5ever passes all tokenizer tests from html5lib-tests, with most tree builder tests passing outside the unimplemented features, and that it does not provide a DOM tree representation.

Official sources

  1. Issues
  2. README
  3. Releases
  4. servo/html5ever on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/servo-html5ever.svg)](https://hysenlabs.com/projects/servo-html5ever)