scraper: CSS selector queries over html5ever in Rust
HTML parsing and querying with CSS selectors
At a glance
- What is it?
- The scraper crate wraps Servo's html5ever and selectors so Rust programs can parse HTML the way a browser does and query it with CSS. It is small, ISC licensed, and deliberately stops at parsing and querying.
- Who is it for?
- Adopt scraper when you need browser-grade HTML parsing inside a Rust program and your queries are expressible as CSS selectors: parse with Html::parse_document, compile selectors once with Selector::parse, and read values through ElementRef. Do not adopt it if you need XPath, a headless browser, or JavaScript execution, because the crate covers parsing and querying only.
- Can I use it commercially?
- Yes. ISC is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 9 days ago.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What scraper solves, and who reaches for it
HTML in the wild is malformed. Tags go unclosed, attributes repeat, and a parser that insists on well-formed XML will reject pages that browsers render without complaint. scraper exists to give Rust programs the same tolerance: the README describes it as providing "an interface to Servo's html5ever and selectors crates, for browser-grade parsing and querying." That sentence is the whole scope. It is a parsing and querying library, not a fetching library, not a headless browser, and not a JavaScript runtime.
The audience follows from that. If you already have HTML bytes, from a file, a test fixture, a response body your own HTTP client fetched, and you want to pull structured values out of them, scraper is the layer you want. It suits CLI tools, one-off extraction scripts, and services that receive HTML and need to normalise it. It does not suit anything that requires the page to execute scripts first, and it does not fetch anything on your behalf.
The mechanism: html5ever parses, selectors matches, ElementRef reads
The data flow has three stages. First, Html::parse_document or Html::parse_fragment turns a string into a tree, using html5ever's tree builder. Second, Selector::parse compiles a CSS selector string into a reusable Selector value; the README's example uses Selector::parse("h1.foo").unwrap(), so a malformed selector surfaces as an error at that point rather than during matching. Third, you call select on the document or on an element with that compiled selector and iterate the results.
What you iterate over is not a bare node. The README shows elements answering value().name(), value().attr("value"), html(), inner_html(), and text(). text() returns an iterator of string pieces, and the README's assertion is exact about the split: for the fragment <h1>Hello, <i>world!</i></h1>, the collected vector is ["Hello, ", "world!"], two entries rather than one concatenated string. If you want a single string you join them yourself. That detail matters more than it looks: it means scraper reports the text nodes as the tree actually holds them, and it does not guess where whitespace should collapse.
Selectors are also scoped. The README's descendant example selects the ul element from the fragment and then calls select on that element, so matching runs against the subtree rather than the whole document. That is the difference between a document-wide query and a query rooted at a node you already found, and it is the pattern you want when a page repeats a block many times.
The crate also exposes a sink for editing. The README's DOM example imports TreeSink from html5ever, collects node ids from a selection, wraps the document in HtmlTreeSink::new(document), calls remove_from_parent for each id, and then calls tree.finish() to get the document back. The assertion shows the removed paragraph gone from the serialised output. So mutation is possible, but it goes through an explicit sink wrapper and a finish step rather than through methods on the document itself.
Installing scraper and running a first extraction
The crate is published on crates.io, and the README points there and to GitHub as the two places it lives. Adding it to a Cargo project is a dependency line. The README's own example for the atomic feature pins version 0.27.0, which is the version string to start from:
[dependencies]
scraper = { version = "0.27.0", features = ["atomic"] }That features line is only needed if you intend to move parsed trees across threads; the README explains the reason in the threading section below. If you do not, drop the features key and take the default.
A first real use is the selection loop the README demonstrates. Parse a fragment, compile a selector, iterate:
use scraper::{Html, Selector};
let html = r#"
<ul>
<li>Foo</li>
<li>Bar</li>
<li>Baz</li>
</ul>
"#;
let fragment = Html::parse_fragment(html);
let selector = Selector::parse("li").unwrap();
for element in fragment.select(&selector) {
assert_eq!("li", element.value().name());
}Running that gives you three iterations, one per list item, and each element reports its tag name as "li". From there the two accessors you will use most are value().attr(...) for attributes and text() for content. The README's attribute example selects input[name="foo"] and asserts that value().attr("value") is Some("bar"), and its text example collects h1.text() into a vector. One practical note the README makes implicitly by using unwrap: Selector::parse returns a Result, so a selector built from user input needs handling rather than unwrapping.
Threading, the atomic feature, and what breaks without it
This is the sharpest constraint in the documentation, and it is easy to miss. The README states that html5ever uses the Tendril type as its reference counted string, that Tendril is thread-local by default and therefore !Send, and that enabling the atomic flag switches to the atomic counting version, which implements Send. In plain terms: by default you cannot move a parsed document to another thread. If your design is a worker pool where one thread parses and others query, the default build will not compile that way.
The fix is the feature flag shown above, scraper = { version = "0.27.0", features = ["atomic"] }. The README does not discuss what atomic reference counting costs in throughput, and it does not say whether every public type becomes Send or only some. That is a gap worth knowing about before you design around it: the documentation tells you the flag exists and what it changes at the Tendril level, and stops there.
A second, quieter limitation is scope. scraper parses and queries. It does not fetch, it does not render, and it does not run JavaScript. A page whose content is assembled client-side will parse into whatever the server actually sent, which may be an empty shell. That is not a defect in the crate; it is the boundary of what browser-grade parsing means. When you cross that boundary you need a different tool, not a different selector.
scraper against a full headless browser
The obvious alternative for HTML extraction in Rust is driving a real browser, for example through a WebDriver client, so the page executes its scripts before you query it. The difference in approach is where the DOM comes from. scraper builds the tree itself from the bytes you hand it, using html5ever, entirely in process. A browser-backed tool builds the tree in a separate rendering engine, after scripts, layout and network activity.
That changes three things. Cost: a browser process is heavy compared with a library call, and you pay it per page. Fidelity: a browser sees what a user sees, including injected content, while scraper sees the response body. Determinism: scraper's output is a pure function of the input string, so a test fixture always parses the same way, whereas a browser-backed run depends on the page's own code and its network calls.
For static HTML, server-rendered pages, and fixtures in a test suite, scraper is the lighter and more predictable choice, and it is the only one of the two that belongs in a unit test. For single-page applications, login flows, or anything gated behind script execution, scraper is the wrong tool and no amount of selector work will fix it. The honest framing is that these are complementary: fetch with one, parse with the other, and only reach for the browser when the content genuinely requires it.
Maintenance, releases and the ISC licence
The repository is not archived, and the last push was on 2026-09-21, which is recent. Releases are tagged rather than continuous: v0.27.0 on 2026-05-11, v0.26.0 on 2026-03-18, v0.25.0 on 2025-12-06. The cadence is a few months between minor versions, and the 0.x prefix means minor bumps can carry breaking changes, so pinning a version in Cargo.toml is the sensible default rather than tracking the latest.
Upgrade cost is the usual Rust story. Because the crate sits on html5ever and selectors, a scraper upgrade can pull in a new html5ever, and the README's DOM example already reaches into html5ever directly for TreeSink. If you use that mutation path, an html5ever bump is your problem as much as scraper's. If you only parse and select, upgrades are more contained. The README does not document a deprecation policy or a migration guide, so the changelog on the release page is the source to read before bumping.
The licence is ISC, a permissive licence, and the repository carries a LICENSE file at the top level. Permissive terms generally mean you can use the crate in closed-source software provided you keep the copyright notice and licence text with it. That is a description of what the licence family typically requires, not legal advice; read the LICENSE file and your own obligations before shipping.
Editorial conclusion
Adopt scraper when you need browser-grade HTML parsing inside a Rust program and your queries are expressible as CSS selectors: parse with Html::parse_document, compile selectors once with Selector::parse, and read values through ElementRef. Do not adopt it if you need XPath, a headless browser, or JavaScript execution, because the crate covers parsing and querying only. Before writing production code, check the current version and feature flags on crates.io, confirm whether you need the atomic feature for cross-thread use, and read the docs.rs page for the exact API surface of the release you pin.
Frequently asked questions
How do I use scraper to select elements from an HTML string?
Parse the string with Html::parse_document or Html::parse_fragment, compile your query with Selector::parse, then iterate fragment.select(&selector). Each result is an element you can read with value().name(), value().attr(...), text(), html() or inner_html().
How do I install scraper in a Rust project?
It is published on crates.io, so it goes in Cargo.toml as a dependency. The README's example pins scraper = { version = "0.27.0", features = ["atomic"] }, with the atomic feature needed only for moving parsed trees across threads.
Can I use scraper across multiple threads?
Not with the default build. The README states html5ever's Tendril string is thread-local and therefore !Send, and that enabling the atomic feature switches to the atomic counting version, which implements Send.
Does scraper fetch web pages for me?
No. The README describes it as HTML parsing and querying with CSS selectors, an interface to html5ever and selectors. Fetching and JavaScript execution are outside what the crate does.
Can scraper modify the DOM, not just read it?
Yes, through a sink. The README's example wraps the document in HtmlTreeSink::new(document), calls remove_from_parent for collected node ids, and then calls tree.finish() to get the document back.
What licence does scraper use?
ISC. The repository lists ISC as the licence and carries a LICENSE file at the top level, alongside Cargo.toml and the scraper directory.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/rust-scraper-scraper)