spider-rs/spider: a Rust crawler that streams pages and launches Chrome only when needed
Get web data for AI agents and LLMs - fast, efficient, and reliable with Rust
At a glance
- What is it?
- Spider is an MIT-licensed Rust crawling engine with bindings for Node.js and Python, a CLI, and an MCP server. Its distinguishing choice is HTTP-first fetching with headless Chrome reserved for pages that need JavaScript, and a Spider Cloud path that reuses the same API.
- Who is it for?
- Adopt Spider if you are building in Rust and want one engine that spans a local script and a distributed worker fleet, or if you want the same crawl code to talk to managed proxies without a rewrite. Do not adopt it if you need a long-stable 1.x API, or if you are not prepared to write Rust for the library path; the Node.js and Python packages exist but the documentation in this repository centers on the Rust crate.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Spider is for, and who should care
Spider targets one job: turning a site into a stream of fetched pages that a program can consume. The README frames the audience in terms of downstream use, naming vector stores for LLM and RAG pipelines, SEO and price monitoring, Markdown, JSON or WARC export, and AI browsing agents. That list is broad, but the underlying requirement is narrow. You have a set of URLs, you want the HTML and status codes, and you do not want to write the link-following, concurrency, retry and browser-launching layers yourself.
The project is Rust-first. The workspace in Cargo.toml lists spider, spider_worker, spider_cli, spider_utils, spider_agent, spider_agent_types, spider_agent_html and spider_mcp as members, so the crawler core, a worker binary, a CLI and an MCP server ship from one repository. Node.js and Python packages are published separately. If your stack is Rust, Spider is a library you embed. If it is not, the CLI and the MCP server are the entry points that do not require you to write Rust at all.
The pitch that deserves scrutiny is scale. The README claims the same engine runs Spider Cloud and that you can move from a local script to managed infrastructure with one config change. The mechanism for that is a SpiderCloudConfig object attached to the Website builder, not a different crawl API. That is a real design commitment, and it is the reason the cloud section appears before the local section in the README.
The HTTP-first, Chrome-on-demand mechanism
The core architectural decision is stated plainly: Spider runs HTTP-first and launches headless Chrome only when a page needs JavaScript. Both paths stream. Pages arrive as they are fetched rather than in a batch at the end, which is why the README's local example subscribes to a channel and prints each page inside the receive loop.
That streaming model shapes the API. You call subscribe with a buffer size, spawn a task that drains the receiver, then call crawl and await it. The crawler discovers links, respects the limits you configured, and stops on its own. Nothing in the example collects results into a vector first, and that is deliberate: for a large crawl, buffering everything before you touch it is the thing you are trying to avoid.
The Chrome path is opt-in twice over. You enable the chrome feature in Cargo.toml, and you call crawl_smart() rather than crawl(). The README describes crawl_smart as trying HTTP first and launching Chrome only on pages that need it. This is the part of the design that has the largest effect on resource use, because a headless browser per concurrent request is expensive, and a hybrid strategy means you pay that cost on the minority of pages that require rendering.
The cloud path follows the same logic one level up. SpiderCloudMode::Smart is documented as routing through proxies first and escalating to the unblocker only on pages that fight back, so bypass cost is incurred per page rather than per crawl. Whether that escalation heuristic matches your targets is something only your own traffic will tell you.
Installing Spider and running a first crawl
The README gives a table of install commands keyed by what you want. For a Rust library it is cargo add spider. For the command-line tool it is cargo install spider_cli. The Node.js package is @spider-rs/spider-rs and the Python package is spider_rs. An MCP server for clients such as Claude and Cursor installs with cargo install spider_mcp.
The smallest working library program uses tokio re-exported from the crate, so you do not add a separate runtime dependency. The README shows this exact shape:
use spider::{tokio, website::Website};
#[tokio::main]
async fn main() {
let mut website = Website::new("https://example.com");
let mut rx = website.subscribe(16);
tokio::spawn(async move {
while let Ok(page) = rx.recv().await {
println!("{} {}", page.status_code, page.get_url());
}
});
website.crawl().await;
website.unsubscribe();
}What you should see is one line per fetched page, each carrying a status code and a URL, printed while the crawl is still running. The 16 passed to subscribe is the channel buffer, so if your consumer is slower than the crawler you will feel backpressure there rather than in memory growth.
Once that runs, the configuration builder is where the crawl becomes useful. Every option has a default, and the README's example sets only what matters for a polite crawl:
let mut website = Website::new("https://example.com")
.with_limit(50) // concurrent requests
.with_depth(10) // how deep to follow links
.with_delay(500) // pause between requests (ms)
.with_respect_robots_txt(true)
.with_subdomains(true)
.with_user_agent(Some("MyBot/1.0"))
.with_stealth(true)
.build()
.unwrap();Note that with_limit is concurrent requests, not a page cap, and with_depth counts link hops, not pages. The full option list lives in the Configuration docs on docs.rs rather than in the README, so budget time for that page before you tune anything.
For JavaScript-heavy targets, the README says to enable the chrome feature and call crawl_smart(). The feature flag goes in your dependency declaration:
[dependencies]
spider = { version = "2", features = ["chrome"] }The repository also ships more than fifty runnable examples under examples/, covering areas such as anti-bot handling, caching, budget control, screenshots and remote Chrome. Those files are the closest thing to a tour of the engine's edges.
Where Spider is the wrong tool
The most concrete limitation visible in the README is version churn. The dependency examples pin spider = "2", the install table tells you to run cargo add spider, and the recent releases jump from v2.48.2 and v2.48.4 in late March to v2.52.2, whose tag is specifically about CLI authentication for Spider Cloud. A release train that adds a cloud authentication path as a headline CLI change is moving quickly, and code written against an early 2.x configuration API may need edits on upgrade. The workspace Cargo.toml also patches the crates.io spider crate to a local path, which is a local-only construct the file itself notes is ignored on publish, but it tells you the tree is actively reshaped.
Second, the README's own framing of the hard part is worth taking literally. It says the hardest part of crawling at scale is not the code but the proxies, headless browsers and anti-bot churn, and that Spider Cloud runs all of that behind the same API. Read as a statement about the library, that is an admission that the open source crawler does not solve anti-bot bypass for you. The with_stealth(true) option exists, and examples/anti_bots.rs is in the repository, but the README does not document what stealth does or how effective it is against any specific defense. If your targets actively block crawlers, the honest reading is that you are buying infrastructure, not a guarantee.
Third, the documentation surface is split. The README covers install, a minimal example and a configuration sample. It does not document error handling, retry semantics, memory behavior on very large crawls, or rollback between versions. For a crawler you intend to run unattended, those are the questions that matter, and they are answered, if at all, in the docs.rs pages and the examples directory rather than in the README.
How Spider differs from a general-purpose scraping framework
The obvious comparison is a Python scraping framework such as Scrapy. The difference is not the feature list; it is where each one puts the concurrency boundary. Scrapy is built around a Twisted reactor and a scheduler that you extend with middlewares and pipelines, and its JavaScript rendering story is an add-on (scrapy-playwright or Splash) that you install and wire up. Spider instead bakes the browser decision into the crawl call itself: crawl_smart() tries HTTP and escalates per page. If you already have a Scrapy deployment with custom middlewares, moving to Spider means rewriting that layer in Rust, and the return is throughput and a single binary rather than a plugin ecosystem.
Against a hosted scraping API, the difference is the opposite direction. A hosted API gives you a URL and returns extracted content, and you own no crawl loop. Spider gives you the loop, the link discovery, the depth and concurrency controls and the streaming channel, and lets you point the same code at either your own infrastructure or Spider Cloud through SpiderCloudConfig. The trade is control against operational burden: with the library you own retries, rate limits and browser processes, and the README is direct that this is the part that gets hard.
Within the Rust ecosystem, Spider's nearest neighbor is a lower-level HTTP client plus an HTML parser that you drive yourself. That path gives you total control and no surprises, at the cost of writing link extraction, deduplication, robots.txt handling, depth limits and the Chrome fallback yourself. Spider's value proposition is precisely that those are already written and exposed as builder methods.
Maintenance, licensing and upgrade cost
The repository is not archived, and the last push was on 2026-09-05, which is recent enough that the project is being worked on. The most recent tagged release listed is v2.52.2 from 2026-03-31, so there is a gap between the last release and the last commit; if you depend on tagged versions rather than the main branch, expect fixes to appear in git before they appear on crates.io.
The licence is MIT, declared in the README badge and in the LICENSE file at the repository root. MIT is permissive: it allows commercial and closed-source use with attribution and without a copyleft obligation on your own code. One thing to check yourself, without treating this as legal advice, is the dependency tree. The workspace patches the external spider_transformations crate onto the local spider crate and mentions keeping the tree on one patched quick-xml, which indicates the project tracks a security-relevant XML parser deliberately. Your own licence and vulnerability review should cover the transitive dependencies your build actually resolves, not just Spider itself.
Upgrade cost is the practical concern. The dependency examples in the README use a floating "2" version, so a cargo update can move you across minor releases without a code change on your side. Given the release cadence, pin an exact version in production and read CHANGELOG.md before bumping. The CLI is a separate crate (spider_cli) with its own release notes, and v2.52.2's note about CLI authentication for Spider Cloud means CLI and library versions can carry different feature sets.
Editorial conclusion
Adopt Spider if you are building in Rust and want one engine that spans a local script and a distributed worker fleet, or if you want the same crawl code to talk to managed proxies without a rewrite. Do not adopt it if you need a long-stable 1.x API, or if you are not prepared to write Rust for the library path; the Node.js and Python packages exist but the documentation in this repository centers on the Rust crate. Before committing, verify three things: that the crate version you pin still exposes the configuration methods your code calls, since the README shows a 2.x line while the workspace patches the crates.io spider crate to a local path, that your JavaScript-heavy targets actually benefit from crawl_smart(), and that your robots.txt and rate-limit policy matches what with_respect_robots_txt and with_delay will do by default.
Frequently asked questions
How do I install the Spider Rust crawler?
The README's install table gives cargo add spider for the Rust library and cargo install spider_cli for the command-line tool. Node.js and Python users install @spider-rs/spider-rs and spider_rs respectively.
Does Spider need headless Chrome to crawl a site?
No. The README states Spider runs HTTP-first and launches headless Chrome only when a page needs JavaScript. To use that path you enable the chrome feature and call crawl_smart() instead of crawl().
Can I use Spider without running my own proxies and browsers?
Yes, through Spider Cloud, which the README says runs proxies, headless browsers and anti-bot handling behind the same API. You attach a SpiderCloudConfig to the Website builder; SpiderCloudMode::Smart routes through proxies first and escalates to the unblocker only on pages that fight back.
What is Spider's licence?
MIT, per the README badge and the LICENSE file at the repository root. That permits commercial and closed-source use with attribution, but you should review the transitive dependencies your own build resolves.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/spider-rs-spider)