# ai.robots.txt: A Maintained Blocklist for AI Crawlers Across Six Web Servers

> ai.robots.txt is a community-maintained repository that publishes a robots.txt file and matching web server configuration snippets for Apache, Nginx, Caddy, HAProxy, and Lighttpd, all pointing at a single source list of AI-related crawler user-agent strings.

**ai-robots-txt/ai.robots.txt** — A list of AI agents and robots to block.

- Repository: https://github.com/ai-robots-txt/ai.robots.txt
- Website: https://github.com/ai-robots-txt/ai.robots.txt/releases.atom
- Stars: 4,155 · Forks: 186
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/ai-robots-txt-ai-robots-txt

## What Problem ai.robots.txt Solves

Site owners who want to prevent AI crawlers from indexing their content face a practical problem: each crawler announces itself with a distinct user-agent string, and the list of active AI crawlers grows faster than any individual can track manually. The ai.robots.txt project maintains a single source of truth for those strings and generates the server-specific configuration files automatically from that source.

The repository targets site owners who use any of the major web servers or proxies: Apache httpd (.htaccess), Nginx (nginx-block-ai-bots.conf), Caddy (Caddyfile), HAProxy (haproxy-block-ai-bots.txt), and Lighttpd (lighttpd-block-ai-bots.conf). Each file is generated from robots.json, the canonical source, via a GitHub Actions workflow. This means a contributor adds or updates a crawler in robots.json, and the derivative files are rebuilt automatically in the next release.

The README notes that the project's crawler data has been sourced in part from Known Agents, a third-party directory that tracks AI crawlers. The list covers crawlers regardless of their stated purpose, meaning training scrapers, search indexers, and summarization bots are all included.

Microsoft Bing presents a specific case the robots.txt file cannot address on its own. The README notes that Bing uses the content it crawls for AI and training, but the standard user-agent block does not cover this use case. The opt-out requires adding a specific meta tag to the head element of each page. The repository links to a dedicated documentation page for the Bing-specific opt-out steps, acknowledging that a complete AI-blocking strategy may require actions beyond the config files this repository provides.

## Six Deployment Formats from One Source File

The repository ships six ready-to-use files, each targeting a different part of the web infrastructure stack.

robots.txt implements RFC 9309, the Robots Exclusion Protocol standard. Placing this file at your domain root communicates blocking intent to crawlers that respect the protocol.

.htaccess configures Apache httpd to return an error page when a listed crawler connects. The README notes that .htaccess has a performance cost compared to main-config directives, and links to the Apache documentation for alternatives.

nginx-block-ai-bots.conf is a snippet to include in any virtual host server block via the include directive.

The Caddyfile provides a Header Regex matcher group. According to the README, the rejection is then handled with:

```
abort @aibots
```

HAProxy integration requires two steps. First, add the file to HAProxy's config directory. Then add these lines to the frontend section:

```
acl ai_robot hdr_sub(user-agent) -i -f /etc/haproxy/haproxy-block-ai-bots.txt
http-request deny if ai_robot
```

The path to haproxy-block-ai-bots.txt in the acl directive must match where the file is placed in your environment.

Lighttpd uses an include directive:

```
include "fragments/lighttpd-block-ai-bots.conf"
```

This can be placed globally or inside a conditional section in lighttpd.conf.

All six files derive from robots.json via a GitHub Actions workflow. Contributors should modify only robots.json, not the generated files, since those are overwritten on each push. This design keeps the crawler list in a single place and ensures all output formats stay in sync without manual editing.

## Following Releases and Contributing New Crawlers

The project ships versioned releases on GitHub, with v1.52 released on 2026-09-07 adding Diffbot-User, OAI-Adsbot, and other agents. Subscribers can track new releases through the RSS/Atom feed at the repository's releases URL, or through Feedly, Inoreader, The Old Reader, Feedbin, or any other feed reader. GitHub users can also subscribe through the Watch button by selecting Custom and then Releases.

Updates must go into robots.json. The GitHub Actions workflow regenerates the derived files from that source. Contributors who want to add a crawler submit a pull request that modifies robots.json, not the generated files directly. The README explicitly prohibits AI-generated contributions.

For contributors who want to run the test suite locally, the Python test dependencies are:

```console
pip install -r requirements.txt
```

With the dependencies installed (beautifulsoup4, lxml, requests), the tests run with:

```console
code/tests.py
```

The repository also includes an .editorconfig file to standardise formatting across editors.

The release process for administrators follows documented steps: navigate to the GitHub new release page, create a new tag as v1.n, write a title in the format v1.n: adds user-agent1, user-agent2, click Generate release notes, and publish. A GitHub Actions workflow then attaches the robots.txt file as a release asset. Because the tagged asset is generated automatically, each release URL contains the robots.txt that corresponds exactly to that version's crawler list, making it safe to pin a specific version URL in a production configuration.

## The Central Limitation: Bots That Ignore the Protocol

The Robots Exclusion Protocol is a voluntary standard. A crawler that chooses to ignore robots.txt will crawl regardless of what the file says. The README acknowledges this directly: it links to a Cloudflare form for reporting abusive crawlers that do not respect robots.txt, and references Cloudflare's hard block feature alongside this list.

User-agent blocking is also bypassable. A crawler operator can change the user-agent string their tool sends. If a known crawler is renamed or rotated, the block list does not automatically detect and block the new string until someone notices and submits a pull request. The table-of-bot-metrics.md file in the repository provides additional context on individual crawlers, but does not document which ones are known to ignore robots.txt.

The .htaccess and web server config files address the bypass risk differently from robots.txt: they enforce the block at the server level regardless of whether the crawler respects the protocol. A request that matches the user-agent pattern receives an error response before any content is served. This is stronger than a robots.txt hint but still relies on the user-agent string being correct.

Sites that need to block crawlers that actively spoof or randomise user-agents are outside the scope of this project.

## Cloudflare as a Network-Layer Alternative

Cloudflare offers AI bot blocking as a feature in its security dashboard. The README links to Cloudflare's documentation on this feature, titled "declaring your AI independence." The key architectural difference is enforcement layer: Cloudflare blocks traffic at the network edge, before a request reaches the origin server, and uses signals beyond the user-agent string (IP reputation, TLS fingerprint, behavioral analysis) to identify bots.

The ai.robots.txt list works at the application layer: a server configuration file or a robots.txt hint. It requires no third-party service, no DNS change, and no traffic routing through an external CDN. For sites that host their own infrastructure and do not use Cloudflare, this repository provides a self-contained solution. For sites already behind Cloudflare, the platform's bot management may be both simpler to maintain and harder to circumvent.

The README suggests combining both approaches: using Cloudflare's hard block for enforcement and reporting crawlers that circumvent it to the form linked in the repository.

## License, Related Projects, and RSL Integration

The repository is MIT licensed. The MIT license permits free use, modification, and redistribution of the configuration files and list data.

The README documents a number of related projects. The Traefik plugin automatically applies the robots.txt rules as middleware, meaning site operators using Traefik can enforce the blocklist without modifying individual server configs. Bot Ledger is a free, static directory of verified AI crawlers with a one-click robots.txt and llms.txt generator that requires no signup. KI-Zugangsindex measures blocking adoption across German .de domains on a fixed panel of 600 domains so the same sites can be re-checked over time. The AI Crawler Census tracks which crawlers the Tranco top 5,000 sites allow or block, with per-domain results published for each run under CC BY 4.0. AI Discovery Radar measures monthly how many websites publish robots.txt, llms.txt, ai.txt, tdmrep.json, and similar files, and whether those files are actually fetchable. Cloudflare's registry of verified bots at radar.cloudflare.com/traffic/verified-bots is mentioned in the README as a companion resource for identifying legitimate crawlers that a site may want to allow rather than block.

The repository also mentions Really Simple Licensing (RSL), a standard that allows site owners to license their content to AI companies inside robots.txt rather than blocking them outright. A WordPress plugin implementing RSL is linked in the README. This is a distinct use case from blocking: RSL is for sites that want to permit AI access under negotiated terms rather than refuse it entirely.

## Conclusion

ai.robots.txt suits site owners who want a maintained, versioned starting point for blocking AI crawlers using the Robots Exclusion Protocol or web server rules. It is not the right solution for blocking bots that actively ignore robots.txt or that randomise their user-agent strings, since the list can only block what it can identify by name. Before deploying it, verify the list against the table-of-bot-metrics.md file in the repository to confirm which crawlers are covered and check whether any entries have changed since the version you pinned.

## FAQ

### Is robots.txt legal?

The Robots Exclusion Protocol is a voluntary technical standard, not a legally binding instrument on its own. Legal enforceability of robots.txt-based crawling restrictions varies by jurisdiction and case law. The ai.robots.txt project implements the standard as defined in RFC 9309 but does not offer legal guidance on enforcement.

### Does robots.txt actually work?

robots.txt works only with crawlers that voluntarily respect the Robots Exclusion Protocol. The README acknowledges this and links to Cloudflare's hard block as a network-level complement for crawlers that ignore the protocol. The web server configuration files in this repository (.htaccess, nginx, HAProxy, Caddy, Lighttpd) enforce blocks at the server level regardless of protocol compliance.

### Can robots.txt be ignored?

Yes. A crawler that does not implement or comply with RFC 9309 will crawl regardless of what robots.txt says. The ai.robots.txt project includes server-level configuration files for Apache, Nginx, Caddy, HAProxy, and Lighttpd that enforce blocks at the connection level, but those still depend on the user-agent string being accurate.

### How to fix blocked by robots.txt error?

If your own crawler or scraper is being blocked by a site using this list, the block reflects the site owner's intent to disallow AI crawlers. The ai.robots.txt project does not have an opt-out mechanism. For legitimate research crawlers, the README notes that contributions can be made to the repository to add context about a crawler's purpose in table-of-bot-metrics.md.

## Sources

- [ai-robots-txt/ai.robots.txt on GitHub](https://github.com/ai-robots-txt/ai.robots.txt)
- [License: MIT](https://github.com/ai-robots-txt/ai.robots.txt/blob/main/LICENSE)
- [Project website](https://github.com/ai-robots-txt/ai.robots.txt/releases.atom)
- [README](https://github.com/ai-robots-txt/ai.robots.txt/blob/main/README.md)
- [Releases](https://github.com/ai-robots-txt/ai.robots.txt/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ai-robots-txt-ai-robots-txt
