AnswerDotAI/llms-txt: the Python tooling behind the /llms.txt proposal
The /llms.txt file, helping language models use your website
At a glance
- What is it?
- The llms-txt repository holds the specification for /llms.txt plus two command line tools, llms_txt2ctx and llms_txt2html. It is a small, Apache-2.0 licensed package whose value depends on whether you treat the proposal as a format to publish or a format to consume.
- Who is it for?
- Adopt the llms-txt package if you already publish or consume /llms.txt files and want a parser that follows the written spec rather than a hand-rolled regex. Skip it if you need a validator with a pass/fail exit code or a generator that writes the file for you, because the repository ships neither.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 5 days ago.
- What is it written in?
- Mainly Jupyter Notebook, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What llms-txt actually solves, and for whom
Web pages are built for people. The README states that an HTML page wraps its information in navigation, ads, and JavaScript, and converting it back into clean text is difficult and imprecise. Context windows remain too small for most sites in their entirety, and every wasted token costs time and money. The proposal in this repository answers that with a single markdown file at a known path, plus clean markdown versions of the pages themselves.
The audience is narrow but real. The README says llms.txt files are used most heavily for software documentation, where coding agents follow them to find API references and tutorials. It also names a business outlining its structure and policies, a personal site answering questions about someone's CV, and a school providing course information. What those have in common is a body of text an agent needs to reach without parsing a rendered page.
The repository itself is not the file format's runtime. It is the specification document, a Python package that parses and renders files that follow it, and a notebook-based build. If you only want to publish an /llms.txt, you may never install this package. If you want to consume one programmatically, the package is the reference implementation.
The mechanism: markdown with a fixed shape, plus link relations
The format is deliberately boring. A file named llms.txt sits at the site root or at any subpath, and covers the URLs under that path. Where more than one file applies, agents should use the most specific one, so /docs/llms.txt wins over /llms.txt for a page under /docs/. The required content is a single H1 naming the project or site. Everything else is optional but ordered: an optional byte-order mark, the H1, a blockquote summary containing key information needed to understand the rest of the file, then zero or more markdown sections that are not headings, then zero or more H2 sections holding file lists.
Each file list entry is a markdown list item with a required hyperlink and an optional colon plus notes. The README gives this mock example:
# Title
> Optional description goes here
Optional details go here
## Section name
- [Link title](https://link_url): Optional link details
## Optional
- [Link title](https://link_url)The section named Optional is a convention, not a keyword the parser enforces: it holds secondary links an agent can skip when a shorter context is needed. The second half of the proposal is the markdown mirror. Pages that agents might need should serve a clean markdown version at the same URL, either with .md appended (page.html.md) or with the extension replaced (page.md), and URLs without a file name append index.html.md or index.md. Clients find these through standard link relations, either as HTML link elements or as an HTTP Link header, and the README notes the header form also works for non-HTML resources and can be added in web server or CDN configuration without touching any pages.
The design choice worth pausing on is that the file is markdown rather than XML or JSON. The README is explicit about why: the files are expected to be read by language models, but the information still follows a specific format that classical parsing techniques such as parsers and regex can handle. That dual audience is the whole trick. A JSON manifest would be easier to validate and harder for a model to skim; markdown is the reverse.
Installing llms-txt and running the two command line tools
The package is published as llms-txt and requires Python 3.10 or newer, with fastcore, httpx and mistletoe as dependencies. The pyproject.toml declares two console scripts, llms_txt2ctx and llms_txt2html, both of which are installed by the package itself.
pip install llms-txtAfter that, llms_txt2ctx is on your path and is the entry point most people will use. It reads a published llms.txt file and turns it into a context document. The README does not spell out the flag set, so check the command's own help output before scripting it.
The second script, llms_txt2html, renders a file to HTML. If you are authoring rather than consuming, the fastest first test is to write a minimal file that satisfies the required parts and point the tool at it. The H1 is the only mandatory section, so a file like the mock example in the README is enough to be structurally valid. What you should see is a parsed structure rather than a wall of text: the H1 as the title, the blockquote as the summary, and the H2 section as a list of links. If your file has no H1, or the blockquote appears after other prose, expect the parse to reflect that ordering problem rather than silently fixing it. The specification fixes the order, and the README does not describe any repair step.
Where llms-txt is the wrong tool
The repository does not ship a validator in the sense most people mean. There is no documented command that takes a URL and exits non-zero when the file is malformed, and no online checker hosted by the project. If your workflow needs a gate in CI that fails a pull request when someone breaks the H1 or moves the blockquote, you are building that gate yourself on top of the parsing library.
Nor is it a generator. The README points elsewhere for that: it says documentation platforms generate one automatically, naming Mintlify's docs, and it notes that all nbdev projects now create .md versions of all pages by default. If your site runs on a CMS whose plugin ecosystem has not caught up, the package gives you parsing, not authoring.
The proposal's own scope is another boundary. This is a proposal, and the README frames it that way throughout. Chrome's Lighthouse audits sites for an llms.txt as part of its agentic browsing checks, and the AI labs publish files for their own developer docs, but none of that makes an /llms.txt a requirement for a site to function. A site with no llms.txt still serves HTML to every crawler that asks. Treating the file as a ranking factor or as something search engines require is a misreading of what the repository claims, and the README makes no such claim.
How llms-txt differs from sitemap.xml and robots.txt
The obvious comparison is sitemap.xml, and the difference is intent rather than format. A sitemap enumerates URLs for a crawler to index. An llms.txt file is a curated reading path: the README says agents are expected to view or search llms.txt to find the information they need, then follow the relevant links, and that the links should point to LLM-friendly content such as the markdown versions of pages. The file stays small enough to fit in context, and the detail lives behind the links, fetched only when needed.
robots.txt is a different axis again. It grants or denies access; llms.txt describes structure for clients that have already been granted access. The README says the file is designed to coexist with existing standards, and the two do different jobs, which is why a site can adopt one without touching the other.
The closer alternative is simply publishing clean markdown pages and letting agents find them. That works, and it is half of what the proposal asks for. What llms.txt adds is the index: one small file that tells an agent which of your markdown pages matter and in what order to consider them, without the agent guessing from a site map or crawling everything. Projects that already generate .md mirrors, as nbdev does, get the second half for free and only need to write the index by hand.
Maintenance, licensing and what upgrading costs
The repository is not archived, and the last push was on 2026-09-04. Release history is uneven: 0.0.4 landed on 2024-09-23, then 0.0.5 and 0.0.6 both on 2026-01-29, roughly an hour apart. That pattern suggests small, bundled changes rather than a steady cadence, and the version number still sits in the 0.0.x range with the classifier Development Status :: 3 - Alpha in pyproject.toml. Plan for the parser's surface to move.
The dependency floor is fastcore>=1.12.9, with httpx and mistletoe unpinned above that. In a locked environment this matters little. In an open one, an httpx major release can reach you without a version bump here, and the README documents no compatibility matrix.
Licensing is Apache-2.0, declared in both the LICENSE file and the pyproject.toml license field. Apache-2.0 includes an express patent grant and requires that you retain notices and state changes when you redistribute. That is a permissive licence and generally safe for commercial use, but the specification text and the code carry the same terms, so if you copy the Format section into your own documentation, keep the attribution. This is a description of the licence terms, not legal advice; run anything unusual past your own counsel.
The repository also declares an nbdev entry point and a chkstyle configuration that skips _modidx.py, which tells you the project is built with nbdev from notebooks in nbs/. If you plan to contribute, that build chain is part of the cost of a patch.
Who should adopt llms-txt, and what to check first
Adopt the package if you run software documentation and want a programmatic path from a published llms.txt into a context document, or if you are building tooling that has to interpret files other people publish. The parsing library exists precisely so you do not reimplement the ordering rules, and the spec text in the README is short enough to read in one sitting before you write a line of code.
Do not adopt it if what you actually want is a validator with a pass/fail exit code, an online checker, or a generator that produces the file from your CMS. The repository provides none of those, and the README points at third-party platforms for generation. Do not adopt it expecting search ranking benefits either; the proposal is about agent access, and nothing in the repository claims otherwise.
Verify three things before you commit. First, confirm your own file has the H1 as the only required section and that the blockquote summary comes before any other prose, because the order is fixed and the tooling does not repair it. Second, check whether your documentation platform already emits .md mirrors, since nbdev and Mintlify do and that removes half the work. Third, read the Changes page linked from the README, which describes what moved between v1 and v2 of the proposal, so you are implementing the current revision rather than a blog post from 2024.
Editorial conclusion
Adopt the llms-txt package if you already publish or consume /llms.txt files and want a parser that follows the written spec rather than a hand-rolled regex. Skip it if you need a validator with a pass/fail exit code or a generator that writes the file for you, because the repository ships neither. Before committing, read the Format section of the README and check that your own file has the required H1 and the ordered blockquote, since the spec is stricter than most published examples suggest.
Frequently asked questions
What is an llms.txt file?
It is a markdown file placed at a site root or subpath that gives language models and agents a short, structured index of the site's content. The README describes it as offering brief background information, guidance, and links to detailed markdown files, with an H1 as the only required section.
Is llms.txt mandatory?
No. The README presents llms.txt as a proposal rather than a requirement, and a site with no such file still serves its HTML as before. Adoption is voluntary, though Chrome's Lighthouse audits sites for one as part of its agentic browsing checks.
How to make llms txt?
Write a markdown file named llms.txt with an H1 naming the site or project, a blockquote summary, and H2 sections holding markdown lists of links to LLM-friendly pages. The repository parses and renders such files but does not generate them; the README points to documentation platforms such as Mintlify for automatic generation.
Where to submit an llms.txt file?
There is no submission step described in the repository. The file is published at the site root as /llms.txt or at any subpath such as /docs/llms.txt, and clients discover it through the describedby link relation or by requesting the path directly.
how to use llms txt
Install the llms-txt package and run the llms_txt2ctx script against a published file to turn it into a context document. Agents are expected to view or search the file, then follow the links it lists to fetch the markdown pages they need.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/answerdotai-llms-txt)