mcpdoc: Auditable llms.txt Retrieval for AI-Powered IDEs
Expose llms-txt to IDEs for development
At a glance
- What is it?
- mcpdoc is an archived Python MCP server from the LangChain team that exposes user-specified llms.txt documentation indexes to IDEs like Cursor, Windsurf, and Claude Code, replacing opaque built-in retrieval with auditable tool calls developers can inspect.
- Who is it for?
- mcpdoc suits developers who need to audit exactly which documentation pages a LangChain or LangGraph question drew on. The domain access controls are a genuine differentiator: a remote llms.txt file permits fetches only from its own host domain, not any URL a prompt might suggest.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- No. The owners have archived the repository on GitHub, so it is read-only and no longer receives changes.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 22, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What mcpdoc solves and who it is for
IDEs like Cursor and Windsurf can consume llms.txt index files natively, and Claude Code can read them through its own tooling. The problem is visibility. When an IDE fetches documentation context internally, the developer cannot easily see which URLs it visited, how much context each page added, or whether retrieval stayed within the intended domain. For developers debugging why an AI gave an unexpected answer about a LangChain API, that opacity is a real obstacle.
The Model Context Protocol (MCP) offers a different path. A developer can register an external MCP server, and every call that server handles appears as a named tool call in the IDE's conversation log. mcpdoc is that MCP server for llms.txt files. It was built by the LangChain team and targets developers using LangChain or LangGraph tooling who want to replace silent documentation retrieval with something they can watch and audit.
The repository was archived by its maintainers, with the last push recorded on 2026-08-20. The project is available from PyPI as the mcpdoc package, version 0.0.10 at the time of archival. It runs on Python 3.10 and later.
How mcpdoc routes llms.txt to the fetch_docs tool
At startup, mcpdoc accepts a list of named llms.txt URLs in the form Label:URL. It does not prefetch or cache these files. Instead it registers two tools with the connected MCP host: list_doc_sources returns the configured names and their llms.txt locations, and fetch_docs accepts a URL from within one of those indexes and returns the Markdown content of that page.
The IDE calls these tools on demand during a conversation. Each call appears in the tool-call history that the IDE exposes to the developer. If an agent calls fetch_docs on a specific LangGraph documentation page, the developer sees the exact URL, the domain it belongs to, and the content retrieved. This lets a team trace an AI answer back to its source page or confirm that a particular reference was never fetched at all.
The two-tool design keeps the surface area small. mcpdoc does not index content, rerank results, or maintain state between sessions. It is a thin fetch layer with a security rule attached to it.
Installing mcpdoc and connecting it to Cursor
The README recommends installing the uv package manager before anything else:
curl -LsSf https://astral.sh/uv/install.sh | shWith uv installed, test the server locally over SSE transport. The following command registers two llms.txt files, starts the server on port 8082, and keeps it running in the foreground:
uvx --from mcpdoc mcpdoc \
--urls "LangGraph:https://langchain-ai.github.io/langgraph/llms.txt" "LangChain:https://python.langchain.com/llms.txt" \
--transport sse \
--port 8082 \
--host localhostThe server starts at http://localhost:8082. To verify the tools before connecting an IDE, the README suggests running the MCP inspector in a separate terminal:
npx @modelcontextprotocol/inspectorConnect the inspector to the running server and call list_doc_sources and fetch_docs manually to confirm the setup works.
For use inside Cursor, open ~/.cursor/mcp.json and add a server entry that uses stdio transport. Stdio is the right choice for IDE integration because it avoids running a persistent local HTTP server:
{
"mcpServers": {
"langgraph-docs-mcp": {
"command": "uvx",
"args": ["--from", "mcpdoc", "mcpdoc",
"--urls", "LangGraph:https://langchain-ai.github.io/langgraph/llms.txt LangChain:https://python.langchain.com/llms.txt",
"--transport", "stdio"]
}
}
}After saving, confirm the server appears as running in Cursor Settings under the MCP tab. The README also suggests adding a User Rule in Cursor Settings / Rules that instructs the agent to call list_doc_sources and then fetch_docs for any LangGraph question, rather than relying on the IDE to invoke the tools automatically.
Domain access controls and what they restrict
mcpdoc does not allow the fetch_docs tool to call arbitrary URLs. The restriction depends on how the llms.txt was supplied at startup. When a remote URL is given, the server automatically adds that URL's host domain to a permitted-domains list and nothing else. The LangGraph llms.txt at langchain-ai.github.io causes langchain-ai.github.io to become the only permitted domain. The fetch_docs tool will refuse to retrieve content from any other domain, even if a URL from that domain appears inside the llms.txt index.
When a local file path is given instead of a remote URL, no domains are added automatically. The developer must pass --allowed-domains followed by one or more domain names. Omitting the flag means every fetch call will be refused.
The wildcard --allowed-domains '*' opens access to all domains. The README explicitly notes this should be used with caution. The domain restriction exists to prevent the tool from being directed at unrelated hosts, which is a risk if a compromised or malformed llms.txt contains external links the developer did not intend to permit.
Where mcpdoc cannot help
mcpdoc can only retrieve content from URLs that appear inside a registered llms.txt file. A documentation site that does not publish an llms.txt has no entry point for this tool. The project cannot crawl arbitrary documentation URLs, process PDFs, ingest private wikis, or connect to databases. If the llms.txt index is incomplete or out of date, mcpdoc will fetch exactly what is listed, without signaling that other relevant pages exist.
There is no caching layer. Every call to fetch_docs is a live HTTP request to the target domain. In environments with restricted outbound network access, or on documentation sites with rate limits, this can cause failures that are not obvious from the IDE's perspective. The tool also has no mechanism for combining content from multiple fetched pages into a single context window; that aggregation is the responsibility of whatever agent or IDE plugin is calling the tools.
The archived status is a practical constraint. Python version requirements, MCP protocol changes, or IDE configuration formats that evolve after 2026-08-20 will not be addressed by new releases.
Comparison with native IDE documentation retrieval
Cursor and Windsurf have their own built-in support for consuming llms.txt files without any external MCP server. The IDE reads the index, decides which pages to fetch based on the conversation, and injects the content into context. From the developer's perspective, this is simpler to set up: no server configuration, no uv installation, no mcp.json edits.
The gap is observability. When the IDE fetches documentation through its internal mechanism, the tool calls are not necessarily surfaced in a way the developer can audit. Which exact pages were fetched, and in what order, may not be visible. mcpdoc makes every fetch a named, logged MCP tool call. For a developer who needs to verify that an AI answer came from the correct version of a documentation page and not from a stale or irrelevant fetch, that log is useful evidence.
The trade-off is complexity and maintenance. mcpdoc adds a process that must be started, configured per IDE, and updated when the MCP protocol or the IDE's configuration format changes. Since the repository is archived, those updates will not come from the upstream project.
Editorial conclusion
mcpdoc suits developers who need to audit exactly which documentation pages a LangChain or LangGraph question drew on. The domain access controls are a genuine differentiator: a remote llms.txt file permits fetches only from its own host domain, not any URL a prompt might suggest. Teams outside the LangChain ecosystem will find it limited because it depends on sites that publish llms.txt, which is not universal. The archived status means no new IDE integrations, transports, or security fixes will appear. Before adopting it, verify that the documentation sites you care about expose an llms.txt file and that the file indexes the specific pages relevant to your work.
Frequently asked questions
Does mcpdoc work with Claude Code in addition to Cursor and Windsurf?
The README describes configuration for Claude Desktop via its JSON config file, using the same stdio server entry as Cursor. Claude Code is also listed as a supported MCP host application in the README's overview section, following the same configuration pattern.
Can I use mcpdoc with a local llms.txt file instead of a remote URL?
Yes, but the server will not permit any fetch domains automatically when given a local file path. The --allowed-domains flag must be passed explicitly, listing each domain the fetch_docs tool should be allowed to reach.
What Python version does mcpdoc require?
The pyproject.toml specifies requires-python >= 3.10, so Python 3.10 or later is needed. The dependencies include httpx, markdownify, mcp[cli] >= 1.4.1, and pyyaml.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/langchain-ai-mcpdoc)