spatie/crawler: a PHP link crawler built on Guzzle promises
https://spatie.be/docs/crawler
At a glance
- What is it?
- spatie/crawler is a PHP package that walks a site's links concurrently using Guzzle promises and can fall back to Chrome via Browsershot for JavaScript-rendered pages. It fits scripted audits and sitemap generation, not large-scale crawling.
- Who is it for?
- Adopt spatie/crawler when you need a scripted, in-process crawl of one site from PHP: link checks, sitemap discovery, or feeding a queue of URLs into a Laravel job. Do not adopt it as a general web-scale crawler, and do not expect it to manage politeness, retries or storage for you.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 54 days ago.
- What is it written in?
- Mainly PHP, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 24, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What spatie/crawler does that a foreach loop over links does not
Fetching a page, extracting its anchors and repeating that for each anchor is easy to write and unpleasant to run. Requests happen one at a time, a slow host stalls the whole run, and every site you point the script at needs its own filtering rules. spatie/crawler packages that loop into a class with a small configuration surface: a starting URL, a depth limit, an internal-only filter, and callbacks that fire as pages complete.
The intended audience is a PHP developer who needs a bounded crawl inside an existing application. Typical uses are checking a site's own links, collecting the URL set of a site you own, or scraping a handful of pages on a schedule. The README describes it as a class to crawl links on a website, and the examples are all short. That framing is honest about the scope: this is a library you call, not a service you deploy.
Guzzle promises, a queue, and where Chrome fits in
The concurrency comes from Guzzle promises. The README states that Guzzle promises are used under the hood to crawl multiple URLs concurrently, which is the mechanism behind the speed difference against a naive loop. Instead of blocking on each response, the crawler keeps several requests in flight and resolves them as they return.
JavaScript rendering is a separate path. The README says the crawler can execute JavaScript and therefore crawl JavaScript-rendered sites, and that Chrome and Puppeteer power that feature through spatie/browsershot. That is a meaningful architectural split: the default HTTP path has no browser dependency, and the rendered path does. A page whose links only exist after client-side rendering will be invisible unless you opt into the browser route, and the browser route costs a Chrome process per render.
Control flow is callback-based. onCrawled receives the URL and a CrawlResponse, and the foundUrls() shortcut returns the collected set instead of invoking your callback. A shouldStopCallback is checked before scheduling each next request, which the README describes as the way to stop a crawl based on external state. That is a cooperative stop, not an interrupt: it takes effect between requests, so it will not abort a request already in flight.
Installing spatie/crawler and running a first crawl
The package installs through Composer, and the README's examples assume the Spatie\Crawler namespace is autoloaded. The README does not spell out the install command, but the package is published on Packagist under the name spatie/crawler, which is the name to require. The README's testing section shows the project's own test command:
composer testAfter that, a first crawl is a few lines. The README gives this example, which prints the status code for every crawled URL:
use Spatie\Crawler\Crawler;
use Spatie\Crawler\CrawlResponse;
Crawler::create('https://example.com')
->onCrawled(function (string $url, CrawlResponse $response) {
echo "{$url}: {$response->status()}\n";
})
->start();Running it against a small site should print one line per URL as responses arrive, with the order depending on how the concurrent requests resolve. If you only want the URL set rather than a per-page callback, the README shows the internalOnly() and depth() chain feeding foundUrls(), which returns the collected URLs without you writing output logic.
Before pointing it at a real host, the fake() method is worth knowing. The README gives an example that maps URLs to HTML strings, so crawl logic can be exercised without making HTTP requests. That is the cheapest way to confirm your filtering and depth rules behave before you generate traffic against someone else's server.
The limits you meet on the second day
Politeness is not built in. Nothing in the README mentions robots.txt handling, crawl-delay, or per-host rate limiting, and the concurrency that makes the package fast is also what will get you blocked if you point it at a host that does not want the traffic. You can slow things down and you can stop on external state, but the coordination is yours to write.
Memory is the other constraint. foundUrls() collects URLs in the process, so a wide crawl of a large site grows the returned array with everything discovered. The README does not document a streaming or persistence mode, which means the onCrawled callback is the escape hatch: write each URL out as it arrives instead of accumulating.
JavaScript rendering has its own failure mode. Because it depends on Chrome through Browsershot, an environment without a working Chrome installation cannot render, and the README does not describe a fallback. The plain HTTP path is unaffected, so the practical rule is to treat rendering as an opt-in capability that your deployment has to support rather than a default.
Finally, this is the wrong tool for crawling at scale across many domains. There is no distributed queue, no shared deduplication across processes, and no built-in storage. Those are features of dedicated crawling systems, and the README makes no claim to them.
spatie/crawler against a general-purpose crawler such as Scrapy
The closest comparison is Scrapy, the Python crawling framework. The difference is not language preference, it is where the crawl lives. Scrapy is a framework you build a project around: it has its own scheduler, its own item pipeline for persisting extracted data, middleware for retries and throttling, and a command-line runner for long jobs. spatie/crawler is a class you instantiate inside an application you already have, and it hands you callbacks.
That makes spatie/crawler a good fit when the crawl is one step in a PHP process: a console command, a queued job, a test. It makes it a poor fit when the crawl is the product. If you need resumable multi-hour crawls, rotating proxies, or a pipeline that writes structured records to a database, Scrapy's architecture already answers those questions and spatie/crawler leaves them to you. The trade is real in both directions: Scrapy asks you to learn a framework, spatie/crawler asks you to write the parts it omits.
Maintenance, upgrades and the MIT licence
The repository is not archived, and the last push was on 2026-08-07, with release 9.4.2 published the same day. That is recent enough that the project is being kept current, and the 9.4.x line shows three releases across July and August 2026.
The repository carries an UPGRADING.md alongside the CHANGELOG.md, which is the signal that major versions have required migration work in the past. The README points to the changelog for what changed recently, and the documentation site is where the full behaviour is described. For a PHP dependency, the upgrade cost is mostly the usual one: read UPGRADING.md before moving a major version, and check whether your callbacks still match the signatures the new version expects.
The licence is MIT, stated in the README and in LICENSE.md. That permits commercial use and modification with attribution, and it means you carry the package's licence text with your distribution. This is a description of what the licence file says, not legal advice; if your organisation has rules about third-party dependencies, the MIT terms are the thing to check against them.
Editorial conclusion
Adopt spatie/crawler when you need a scripted, in-process crawl of one site from PHP: link checks, sitemap discovery, or feeding a queue of URLs into a Laravel job. Do not adopt it as a general web-scale crawler, and do not expect it to manage politeness, retries or storage for you. Before writing code, read the documentation site at spatie.be/docs/crawler, because the README only shows the entry points and the details of depth, filtering and JavaScript rendering live there. Verify the current release on Packagist against the 9.4.x line, and check whether the JavaScript option's Browsershot requirement fits your deployment, since that pulls in a Chrome dependency the plain HTTP path does not need.
Frequently asked questions
How do I install spatie/crawler?
It is a Composer package published on Packagist as spatie/crawler, so it installs with composer require spatie/crawler. The README's examples then assume the Spatie\Crawler namespace is autoloaded.
Can spatie/crawler crawl JavaScript-rendered sites?
Yes. The README states the crawler can execute JavaScript and therefore crawl JavaScript-rendered sites, using Chrome and Puppeteer through spatie/browsershot. That path requires a working Chrome installation, unlike the default HTTP path.
How do I stop a spatie/crawler crawl partway through?
Register a shouldStopCallback, which the README says receives the current crawler instance and is checked before scheduling each next request. Because it is checked between requests, it will not abort a request already in flight.
Can I test spatie/crawler without making real HTTP requests?
Yes. The README shows a fake() method that takes a map of URLs to HTML strings, so crawl logic can be exercised without network calls.
What licence does spatie/crawler use?
The README and LICENSE.md state the MIT licence, which permits commercial use and modification with attribution.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/spatie-crawler)