# google/robotstxt: Google's C++ robots.txt parser as a library

> It is the parser and matcher Googlebot uses, packaged as a C++14 library with a small command line tool. Adopt it if you need Google's exact reading of REP rules; it will not normalise your URIs for you.

**google/robotstxt** — The repository contains Google's robots.txt parser and matcher as a C++ library (compliant to C++11).

- Repository: https://github.com/google/robotstxt
- Stars: 3,474 · Forks: 257
- Language: C++
- License: Apache-2.0
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/google-robotstxt

## The parsing divergence google/robotstxt was written to end

The Robots Exclusion Protocol spent roughly 25 years as a de-facto standard rather than a formal one. The README is blunt about the consequence: different implementers parse robots.txt slightly differently, which produces confusion. A rule one crawler honours another may ignore. That is the problem this repository addresses, and it addresses it by publishing the parser Google itself uses.

The README describes the code as slightly modified production code used by Googlebot, with some internal headers and equivalent symbols swapped out. So the target audience is narrow and technical: people building crawlers, SEO tooling, or compliance checks who want their interpretation of a robots.txt file to line up with Google's. Webmasters get a secondary benefit, a small binary for testing one URL and one user-agent against a local file.

## How the parser and matcher split the work

The repository layout shows two implementation files, robots.cc and robots.h, with robots_main.cc as the command line entry point and robots_test.cc as the test suite. A WebAssembly build target exists as robots_wasm.cc, and reporting_robots.cc handles reporting. That is the whole library: parsing and matching, nothing around it.

The README names the two functions that do the matching, AllowedByRobots and OneAgentAllowedByRobots, and states that parsing is done exactly as in the production version of Googlebot, including how percent codes and unicode characters inside patterns are handled. The boundary is explicit. The library does not perform full normalisation of the URI you pass in. You must supply a URI in the format specified by RFC3986, and only then will matching follow the REP specification.

That constraint is the design decision worth pausing on. A library that normalised URIs would hide bugs in the caller's URL handling but would also pick a canonical form on the caller's behalf. Google chose not to. If your URLs arrive from a redirect chain, a query string builder, or user input, the correctness of the answer depends on your code, not on this one.

## Building google/robotstxt with Bazel or CMake

The README lists the prerequisites: a platform such as Windows, macOS or Linux, a C++ compiler supporting at least C++14, Git, and Bazel if you follow the documented path. Bazel is described as the official build system; CMake is the community-supported alternative. Note the version discrepancy: the repository description says C++11 while the README says C++14, and the README is the one that names the compiler requirement.

The Bazel flow clones the repository, runs the test target, builds the main binary, and then runs it against a local robots.txt, a user-agent and a URL.

```bash
git clone https://github.com/google/robotstxt.git robotstxt
cd robotstxt/
bazel test :robots_test
bazel build :robots_main
bazel run robots_main -- ~/local/path/to/robots.txt YourBot https://example.com/url
```

The README shows the expected output line: user-agent 'YourBot' with URI 'https://example.com/url': ALLOWED. Exit codes carry the answer too, 0 for ALLOWED and 1 for DISALLOWED, which makes the binary usable from a shell script or a CI step without parsing stdout.

The CMake path produces an equivalent binary under a different name, robots rather than robots_main.

```bash
mkdir c-build && cd c-build
cmake .. -DROBOTS_BUILD_TESTS=ON
make
make test
robots ~/local/path/to/robots.txt YourBot https://example.com/url
```

The test target should pass, and the binary prints the same ALLOWED or DISALLOWED verdict. One detail the README calls out: if the robots file is empty, the parser additionally prints a notice that the file is empty so all user-agents are allowed. That notice goes to the output stream, so a script that reads the verdict line should not assume it is the only line printed.

## What google/robotstxt deliberately leaves to the caller

The most useful limitation is stated plainly in the README's notes rather than buried. The library and its binary do not handle implementation logic a crawler might apply outside parsing and matching. The example given is Googlebot-Image respecting rules specified for User-agent: Googlebot when the robots.txt file does not define Googlebot-Image explicitly.

That matters because it is exactly the kind of behaviour people assume a Google library would replicate. It will not. If you are building a crawler that runs several user-agents against one file, the group inheritance logic is yours to write on top of the matcher.

The URI normalisation gap is the second limitation and arguably the sharper one. Passing a URI that is not RFC3986-conformant does not raise an error the README documents; the matching simply will not be done according to the REP specification. That is a silent failure mode. A percent-encoded path, a relative reference, or a URL with an unusual case in the scheme can all produce an answer that looks authoritative and is not.

There is also no published API stability statement in the README, so treat the header as something to pin to a tagged release rather than track from the default branch.

## Where a full crawler stack differs from a parsing library

The obvious alternative is not another robots.txt parser but a crawler framework that already contains one. Scrapy, for instance, exposes a robotstxt_obey setting that turns robots.txt compliance on for a spider. The difference in approach is not quality, it is scope. Scrapy hands you download scheduling, middleware, request deduplication and robots.txt handling as one integrated system, and the robots.txt behaviour is whatever that framework implements.

google/robotstxt gives you one thing: the parse and match step, with the claim that it matches Googlebot. If your question is "will Google crawl this URL", that claim is the whole point and a framework's internal parser does not answer it. If your question is "how do I crawl a site politely at scale", the library is the wrong shape of tool, because you would still need to write everything around it.

A middle option exists in the repository itself: the WebAssembly build target, robots_wasm.cc, suggests the matching logic can be compiled for a browser or edge context where a full C++ toolchain is not available. The README does not document how to build or call that target, so treat it as a lead to investigate in the source rather than a supported path.

## Licence and the cost of keeping up

The library is licensed under the Apache License, version 2.0, with the LICENSE file at the repository root. Apache-2.0 is permissive and includes an express patent grant, which is the usual reason projects pick it over MIT for code that may touch patented techniques. That is a general property of the licence, not a statement about this codebase. If your organisation has rules about which licences may be linked into a shipped product, the identifier to check is Apache-2.0.

The upgrade cost is low by construction. The public surface is small, essentially the two matching functions named in the README plus the header, and the library has no dependencies beyond a C++14 compiler and, for the documented build, Bazel. There is one tagged stable release, v1.0.0, and the last push to the default branch was on 2026-04-01. That is under six months before today, so the repository is not dormant, but it is also a small, stable piece of code rather than a fast-moving one. Pin to the tag and re-check when a new one appears.

## Conclusion

Adopt google/robotstxt when you need the same parsing and matching behaviour Googlebot applies, and you can feed it URIs that already conform to RFC3986. Do not adopt it if you want a crawler framework, a URL canonicaliser, or crawler-level rules such as Googlebot-Image inheriting Googlebot's group, because the README states the library does not handle that logic. Before committing, build the included binary and check one of your own robots.txt files against a real URL and user-agent, then confirm the exit code matches what you expect.

## FAQ

### What is google/robotstxt used for?

It is Google's robots.txt parser and matcher released as a C++ library, described in the README as slightly modified production code used by Googlebot. It is meant for developers building tools that reflect Google's robots.txt parsing and matching, plus a small binary for testing one URL and user-agent against a local file.

### How do I build and run google/robotstxt?

The README documents cloning the repository and using Bazel, the official build system, with bazel test :robots_test, bazel build :robots_main and then bazel run robots_main with a robots.txt path, a user-agent and a URL. CMake is supported as the community build system and produces a binary named robots.

### Does google/robotstxt normalise the URLs I pass to it?

No. The README states the library will not perform full normalisation of URI parameters, and that the caller must ensure the URI follows the RFC3986 format for matching to be done according to the REP specification.

### Does google/robotstxt handle crawler rules like Googlebot-Image inheriting Googlebot's rules?

No. The README says the library and its binary do not handle implementation logic a crawler might apply outside parsing and matching, and gives Googlebot-Image respecting rules set for User-agent: Googlebot as an example of what is left out.

### What exit codes does the google/robotstxt binary return?

The README lists 0 for ALLOWED and 1 for DISALLOWED. If the robots file is empty, the parser also prints a notice that all user-agents are allowed.

### What licence is google/robotstxt released under?

The README states the robots.txt parser and matcher C++ library is licensed under the terms of the Apache license, with the LICENSE file at the repository root carrying the full text.

## Sources

- [google/robotstxt on GitHub](https://github.com/google/robotstxt)
- [Issues](https://github.com/google/robotstxt/issues)
- [License: Apache-2.0](https://github.com/google/robotstxt/blob/master/LICENSE)
- [README](https://github.com/google/robotstxt/blob/master/README.md)
- [Releases](https://github.com/google/robotstxt/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/google-robotstxt
