# llms-txt-hub: a directory of llms.txt implementations and the llmstxt-cli that feeds them to coding agents

> The repository is two things at once: a curated directory of sites that publish llms.txt files, and a CLI that installs those docs as skills into AI coding agents. The directory is the more mature half.

**thedaviddias/llms-txt-hub** — 🤖 The largest directory for AI-ready documentation and tools implementing the proposed llms.txt standard

- Repository: https://github.com/thedaviddias/llms-txt-hub
- Website: https://llmstxthub.com
- Stars: 906 · Forks: 757
- Language: TypeScript
- License: NOASSERTION
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/thedaviddias-llms-txt-hub

## The gap llms-txt-hub fills between a spec and its adopters

The llms.txt idea is a convention, not a runtime. A site publishes a plain text file at /llms.txt that tells an LLM how to read its documentation and what it is allowed to touch. The README frames the purpose plainly: the file is "a standardized way to provide information about how LLM-powered tools and services should interact with your documentation and codebase." What the convention lacks is an index. If you want to know who actually ships one, what a real file looks like, or which of your tools already consume them, there is no registry to query. This repository is that registry, and it is also a set of clients for it.

The audience splits cleanly. One group is developers and technical writers who want examples before writing their own file, or who want to confirm that a vendor they depend on publishes one. The other group uses AI coding agents and wants a project's documentation pulled into the agent rather than retyped. The first group reads the website; the second installs the CLI. Both live in the same monorepo, which is worth knowing before you clone it, because the root package.json is private and orchestrates everything through Turborepo.

## How the directory is assembled: generated lists, hand-written entries

The README's project list is not maintained by hand. It sits between two HTML comments, LLMS-LIST:START and LLMS-LIST:END, and the comment itself says not to remove or modify that section. Entries are generated from repository data. The root package.json exposes the machinery: generate-websites runs tsx scripts/generate-websites.ts, generate-search-index runs node scripts/search-index-generator.cjs, and generate-llms runs pnpm --filter generator start. The build script chains the search index generator before turbo build, so the index is rebuilt on every build rather than committed stale.

Each entry carries a favicon pulled from Google's favicon service, a name, a link, a one-line description, and one or more links to the published files. Some sites list both /llms.txt and /llms-full.txt, which is the practical distinction the directory makes visible: the short file is a map, the full one carries the content. Entries are grouped by primary category (AI & ML, Developer Tools, Data & Analytics, Integration & Automation, Infrastructure & Cloud, Security & Identity) and by secondary categories such as agency, e-commerce, education, and media. That second axis is unusual and it shows in the data: the agency section includes law firms, branding studios, and immigration assessment services, not just software vendors.

The quality control is scripted. check:websites runs tsx scripts/validate-websites.ts, and a warnings mode exists behind the --warn-descriptions flag. There is a separate check:frontmatter script and a lychee link checker configured through lychee.toml. None of these grade a file against the proposed standard. They check that entries are well formed and that links resolve. Treat the directory as a catalogue, not a certification.

## Installing llmstxt-cli and installing a project's docs as a skill

The CLI is the part you run. It is published to npm as llmstxt-cli, currently at 0.4.1, and the README describes it as installing llms.txt docs as skills into 35+ AI coding agents. The package lives at packages/cli, and its own README is the authoritative source for flags and the supported agent list; the root README only links to it.

The repository's own scripts are the only commands documented in the README and package.json, so the honest first step is to read packages/cli/README.md before running anything. For the directory side, the root package.json defines the workspace commands you would run after cloning.

```bash
pnpm install
pnpm dev
```

The dev script is defined as cross-env FORCE_COLOR=1 turbo dev --parallel, so you should see the workspace apps start in parallel. The root package.json is private, so nothing here is published as a package.

To rebuild the generated lists after editing data files, the build script chains the search index generator before the Turborepo build.

```bash
pnpm generate-search-index
pnpm build
```

Running the monorepo locally also expects environment variables copied from .env.example into .env.local. The example file lists FLAGS_SECRET, GITHUB_TOKEN, GITHUB_CLIENT_ID, GITHUB_CLIENT_SECRET, NEXT_PUBLIC_SENTRY_DSN, SENTRY_AUTH_TOKEN, SENTRY_ORG, and SENTRY_PROJECT. The comments in that file say the GitHub token needs the repo, user, and read:org scopes. None of this is needed to browse the hosted directory at llmstxthub.com.

## What the directory does not do, and where the CLI stops being the right tool

The name invites a wrong expectation. This is not a validator. Nothing in the described scripts parses an arbitrary llms.txt and reports whether it conforms to the proposed standard. check:websites validates the repository's own website entries, check:frontmatter validates frontmatter on content files, and lychee checks that links resolve. If your goal is to grade a file you just wrote, you will not find that here, and the README does not claim otherwise.

The CLI has a narrower failure mode: it is only useful if it targets an agent you actually use. The README states 35+ agents, which is a claim about breadth, not about any particular integration being first class. If your editor is not in the list that packages/cli/README.md documents, the tool does nothing for you. The same applies if your team's documentation is behind authentication, since the input is a URL to a published file.

There is also a coverage problem inherent to any directory. Sites appear because they were submitted and passed the link checks, not because they were audited for accuracy. A listed /llms.txt may be thin, out of date, or describe a product differently from the site's own pages. The directory tells you a file exists and where it is. It cannot tell you the file is good.

## llms.txt versus llms-full.txt, and how this repo compares to a generator

The clearest alternative in the same space is a generator. The README lists one directly: LLMs.txt Generator, described as generating llms.txt from a sitemap, crawled pages, and AI, hosted at llmstxtgenerator.co. The difference in approach is the direction of travel. A generator starts from your site and produces a file, which is what you want when you have nothing published yet. This repository starts from files that already exist and indexes them, which is what you want when you are deciding whether to publish one, or when you need to pull someone else's into your agent.

That distinction also maps onto the llms.txt versus llms-full.txt question. The directory shows both forms side by side on entries like eaucube.com and miraiminds.co, and the pattern is consistent: the short file is a curated map of links, the full file carries the content itself. A generator that crawls pages tends to produce the second kind. An index like this one is useful precisely because it shows you which sites chose which.

For the narrower job of checking whether a site publishes a file at all, the README points to a Chrome extension, LLMs.txt Checker, and there are VS Code, Raycast, and MCP Explorer clients listed alongside it. Those are single-purpose clients over the same directory. If you only need to know whether a URL has a file, the extension is less setup than the CLI.

## Maintenance, licence, and what the monorepo costs to keep running

The repository is not archived, and the last push was on 2026-09-10, which is recent. The release history is lopsided: llmstxt-cli@0.3.0, 0.4.0, and 0.4.1 were all published on 2026-02-13, within hours of each other. The web application is versioned 1.0.0 in the root package.json, but that root package is private and the version field there is not a published artifact. The practical reading is that the CLI carries the versioned release cadence and the directory is updated through commits to data files.

The licence file is present as LICENCE at the repository root, but the repository metadata reports the licence as NOASSERTION, meaning GitHub could not map the file to a recognised identifier. If you intend to reuse the data or the CLI in a commercial product, read LICENCE yourself and, where the terms matter, get advice from someone qualified to give it. Nothing here tells you the terms.

For a self-hosted instance, the recurring cost is not the code, it is the credentials. The .env.example requires a GitHub token with repo, user, and read:org scopes, plus OAuth client credentials, plus a Sentry DSN and auth token, plus a feature flags secret. That is a non-trivial set of secrets to hold for a documentation index, and the token scopes are broader than a read-only directory would suggest. The root scripts also assume pnpm and Turborepo, with postinstall running manypkg fix, so the local toolchain is opinionated.

## Conclusion

Adopt the directory if you want to see who publishes llms.txt and what an llms-full.txt variant looks like next to a plain llms.txt; it is a browsable index, not a validator. Adopt llmstxt-cli if you use one of the coding agents it targets and want a project's llms.txt materialised as a skill rather than pasted into a prompt. Skip it if you need a checker that grades your file against the proposed standard, because the repository ships no such scoring logic. Before you build on either half, open packages/cli/README.md and confirm the agent list and the install path still match your tool, since the CLI is at 0.4.1 while the web app is versioned 1.0.0 in package.json.

## FAQ

### What is llms.txt used for in the llms-txt-hub directory?

The README describes it as a standardized way to tell LLM-powered tools and services how to interact with your documentation and codebase, and the directory lists sites that publish one at /llms.txt. The stated goals include guiding AI models on interpreting your docs, improving the accuracy of AI answers about your project, and setting boundaries for AI interaction with your content.

### Is llms.txt mandatory?

The repository treats it as a proposed standard and a convention, not a requirement. It is a file a site chooses to publish at /llms.txt, and the directory exists to catalogue the sites that do.

### How do I check whether a website has an llms.txt file?

The README lists a Chrome extension called LLMs.txt Checker for exactly this, along with VS Code, Raycast, and MCP Explorer clients that search the same directory. The directory itself also links each entry's llms.txt and, where present, llms-full.txt.

## Sources

- [Issues](https://github.com/thedaviddias/llms-txt-hub/issues)
- [Project website](https://llmstxthub.com)
- [README](https://github.com/thedaviddias/llms-txt-hub/blob/main/README.md)
- [Releases](https://github.com/thedaviddias/llms-txt-hub/releases)
- [thedaviddias/llms-txt-hub on GitHub](https://github.com/thedaviddias/llms-txt-hub)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/thedaviddias-llms-txt-hub
