Library / SDK
JayBizzle/Crawler-Detect avatar
JayBizzle/Crawler-Detect

Crawler-Detect: PHP bot detection from the User-Agent and From headers

đź•· CrawlerDetect is a PHP class for detecting bots/crawlers/spiders via the user agent

2,407 stars280 forksPHPMIT

At a glance

What is it?
CrawlerDetect is a small MIT-licensed PHP class that matches request headers against a regex fixture list of known crawlers. It is a good fit for logging, analytics filtering and light bot mitigation, and the wrong fit for identifying a crawler that lies about its identity.
Who is it for?
Adopt Crawler-Detect if you run PHP and need a cheap, dependency-light way to label known crawlers in logs, analytics or a first-pass filter, and you can accept that a spoofed User-Agent defeats it. Do not adopt it as your only defence against scrapers or credential-stuffing traffic, and do not expect it to identify a bot that sends a genuine browser User-Agent with no self-identifying header.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 24 days ago.
What is it written in?
Mainly PHP, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem Crawler-Detect solves, and who actually needs it

Every PHP site eventually has to answer one question about an incoming request: is this a person or a program? The README frames the library narrowly around that question. CrawlerDetect looks at the User-Agent and HTTP_FROM headers and tells you whether the request matches a known bot. It is not a firewall, not a rate limiter and not a behavioural analyser. It is a lookup.

The audience is therefore specific. If you maintain a PHP application and you want to keep Googlebot out of your page-view counters, tag crawler traffic in your access logs, or serve a different response to known spiders, this library gives you a boolean and the name of the matched bot. The README describes it as recognising thousands of user agents, and the maintenance model is a fixture file plus pull requests, so the coverage grows with the community rather than with a vendor.

Where it stops being the right tool is equally clear. A scraper that copies a Chrome User-Agent and sends no other identifying header looks exactly like Chrome to this library. Detection here is signature matching, not proof of humanity.

How the matching works: fixtures, regexes and the two headers

The repository layout tells most of the story. src/ holds the class, src/Fixtures/Crawlers.php holds the pattern data, tests/data/user_agent/crawlers.txt holds the user agent strings used in tests, and export.php regenerates the raw/Crawlers.json and raw/Crawlers.txt files after a merge. A contributor adds a regex to the $data array in the fixture file and a failing user agent string to the test data file. That is the whole contribution loop.

The part worth reading twice is the header handling. With no constructor arguments, the class reads $_SERVER. You can also hand it a header collection, and the README states that both real header names such as User-Agent and PHP's SAPI names such as HTTP_USER_AGENT are understood, with values as strings or arrays of strings. The reason given is concrete: some crawlers, Googlebot named specifically, send a genuine browser User-Agent and identify themselves in another header such as From or Sec-CH-UA. Passing the full header set lets the class check all of them. Passing a single User-Agent string can only ever check one.

That design choice is the difference between this library and a one-line preg_match in your controller. It also means the constructor signature matters more than it looks: if you pass the wrong shape, you are silently back to single-header matching.

Installing Crawler-Detect and running a first check

Installation is a single Composer command. The package name is jaybizzle/crawler-detect.

bash
composer require jaybizzle/crawler-detect

The README's first example instantiates the class with no arguments, which means it reads $_SERVER, and then calls isCrawler(). The same call also accepts a user agent string if you want to test a value you already have.

php
use Jaybizzle\CrawlerDetect\CrawlerDetect;

$CrawlerDetect = new CrawlerDetect;

if ($CrawlerDetect->isCrawler()) {
    // true if a crawler user agent was detected
}

if ($CrawlerDetect->isCrawler('Mozilla/5.0 (compatible; Sosospider/2.0; +http://help.soso.com/webspider.htm)')) {
    // true if a crawler user agent was detected
}

echo $CrawlerDetect->getMatches();

getMatches() returns the name of the bot that matched, or nothing if there was no match. That is the value you want in a log column, because "a crawler hit this page" is far less useful than "Sosospider hit this page".

If your request object is not $_SERVER, pass the headers in. The README shows three shapes: a PSR-7 request, Symfony's HeaderBag, and Swoole. The advice attached to those examples is to prefer the full header set over isCrawler($request->getHeaderLine('User-Agent')).

php
// PSR-7 (Slim, Mezzio, Laminas, League)
$CrawlerDetect = new CrawlerDetect($request->getHeaders());

// Symfony HttpFoundation
$CrawlerDetect = new CrawlerDetect($request->headers->all());

// Swoole
$CrawlerDetect = new CrawlerDetect($request->header);

if ($CrawlerDetect->isCrawler()) {
    // ...
}

What you should see is a boolean from isCrawler() and, when it is true, a bot name from getMatches(). If isCrawler() is false for traffic you know is a crawler, the fixture list does not cover that user agent yet, and the README's contribution path is to open an issue with the user agent string.

The honest limitation: header spoofing and the single-string shortcut

The library's own documentation makes the strongest argument against using it alone. The note about Googlebot sending a genuine browser User-Agent and identifying itself in a separate header exists because header-based detection is only as good as the sender's honesty. A well-behaved crawler announces itself. A hostile one does not, and Crawler-Detect has no mechanism to tell the difference, because it never verifies the claim. There is no reverse DNS check, no IP range validation and no behavioural signal in the README.

There is a second, quieter failure mode. The README explicitly warns against isCrawler($request->getHeaderLine('User-Agent')) in favour of passing the whole header collection. That warning is easy to skip, and the shortcut is the natural thing to write when you are already holding a PSR-7 request. Taking the shortcut narrows detection to one header and reintroduces exactly the Googlebot case the full-header path was designed to catch.

The third limitation is coverage. Detection is a regex list. A crawler that is not in src/Fixtures/Crawlers.php is not detected until someone adds it, and the README's process for that is a pull request or an issue. For a fast-moving scraper, that lag is real. If your problem is abuse rather than classification, this library is the wrong layer of the stack.

Crawler-Detect compared with MobileDetect and its language ports

The README credits MobileDetect as the basis for parts of the library, and the two projects share a shape: a fixture-driven PHP detector that reads request headers and returns a match. The difference is the question they answer. MobileDetect classifies the device or platform behind a request, which is what you want when the decision is which template or asset set to serve. CrawlerDetect classifies whether the caller is a program, which is what you want when the decision is whether to count the visit, log it differently or block it. They are complementary, not substitutes, and a site that cares about both will end up calling each for different reasons.

The other comparison the repository itself offers is the port list. CrawlerDetect has been ported to Laravel, Symfony, Yii2, Node.js, Python, the JVM, .NET, Ruby and Go, each maintained separately. That matters if your stack is not PHP: the Go port is crawlerdetect, the Node port is es6-crawler-detect, and the Python port is crawlerdetect. The trade-off is that a port's fixture list drifts from the upstream one unless its maintainer keeps pulling changes, and the README gives no synchronisation mechanism. If you need the newest patterns, the PHP original is the reference implementation.

Maintenance, upgrade cost and the MIT licence

The repository is not archived, and the last push was on 2026-09-06. Releases are frequent and small: v1.4.1 on 2026-07-10, v1.4.0 on 2026-06-11, v1.3.11 on 2026-05-10. The pattern suggests incremental fixture updates rather than structural change, which is the cheap kind of upgrade. A minor bump is likely to mean new patterns, and the risk of a breaking API change between patch releases is low given that the public surface is a constructor, isCrawler() and getMatches().

The repository carries a phpstan.neon.dist, a phpunit.xml, a .php-cs-fixer.dist.php and a .coveralls.yml, so static analysis, tests and coverage are part of the project's own workflow. That does not tell you anything about your integration, but it does mean a fixture pull request is expected to arrive with a test case.

The licence is MIT, stated in the README and present as a LICENSE file at the repository root. MIT permits commercial use, modification and redistribution provided the copyright notice and permission notice are retained, but the practical question for an adopter is not the licence text. It is whether the fixture list, which is the library's real value, is something you are comfortable depending on a community pull request process to keep current. That is a maintenance judgement, not a legal one, and it is worth making explicitly before you build a blocking rule on top of this.

Editorial conclusion

Adopt Crawler-Detect if you run PHP and need a cheap, dependency-light way to label known crawlers in logs, analytics or a first-pass filter, and you can accept that a spoofed User-Agent defeats it. Do not adopt it as your only defence against scrapers or credential-stuffing traffic, and do not expect it to identify a bot that sends a genuine browser User-Agent with no self-identifying header. Before wiring it in, verify which headers your framework actually exposes, confirm that getMatches() returns the bot name you want to store, and decide what happens to a request that matches nothing.

Frequently asked questions

How do I detect the Google crawler with Crawler-Detect?

Pass the full header collection rather than a single User-Agent string. The README states that Googlebot in particular sends a genuine browser User-Agent and identifies itself in another header such as From or Sec-CH-UA, so a single-string check can miss it.

Is Crawler-Detect a complete bot detection solution?

No. It matches the User-Agent and HTTP_FROM headers against a regex fixture list, so it identifies crawlers that announce themselves and cannot identify one that sends a genuine browser User-Agent with no self-identifying header.

How do I install Crawler-Detect in a PHP project?

Run composer require jaybizzle/crawler-detect. The README then instantiates Jaybizzle\CrawlerDetect\CrawlerDetect with no arguments to read $_SERVER, or with a header collection from a PSR-7 request, Symfony's HeaderBag, or Swoole.

What does getMatches() return in Crawler-Detect?

It returns the name of the bot that matched, if any. The README shows it echoed after an isCrawler() check, which makes it suitable for storing the matched bot name alongside a log entry.

Can I use Crawler-Detect outside PHP?

The README lists ports for Laravel, Symfony, Yii2, Node.js, Python, the JVM, .NET, Ruby and Go, each maintained as a separate project. The README does not describe how a port stays in sync with the upstream fixture list.

Official sources

  1. JayBizzle/Crawler-Detect on GitHub
  2. License: MIT
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/jaybizzle-crawler-detect.svg)](https://hysenlabs.com/projects/jaybizzle-crawler-detect)