Ref MCP: Token-Efficient Documentation Search for AI Coding Agents
Helping coding agents never make mistakes working with public or private libraries without wasting the context window.
At a glance
- What is it?
- Ref MCP is a Model Context Protocol server that gives AI coding agents search and read access to technical documentation, designed to minimize context window usage by filtering repeated results and returning only the most relevant section of each page rather than full documents.
- Who is it for?
- Ref MCP is the right choice for engineering teams using Claude Code, Cursor, or other MCP-compatible coding agents who are spending significant token budget on documentation lookups that pull in large, mostly irrelevant pages. The streamable HTTP setup requires signing up for an API key at ref.tools; the local stdio server in this repository is the legacy path.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
The Problem: Documentation Context Filling the Context Window
AI coding agents that need to look up API documentation or library behavior typically fetch a documentation page and include it in the context window. A standard web page for a large framework can run 20,000 or more tokens, most of which are irrelevant to the specific question being answered. The README describes this as a documented phenomenon: models perform worse as context grows, and unnecessary context is also expensive at API pricing.
Ref MCP addresses this by acting as a search and retrieval layer that sits between the agent and documentation sources. Instead of fetching a full page, the agent calls ref_search_documentation with a natural-language query, then calls ref_read_url on a specific section identified from the search results. The server uses the session's search history to trim the returned content to the most relevant portion of the target page, typically around 5,000 tokens rather than the full document.
How Ref Minimizes Tokens Through Session-Based Filtering
Ref tracks the search trajectory within a session and applies two token-reduction techniques. First, for repeated similar searches within a session, it never returns the same result twice. This means an agent that refines its search query gets new results rather than the same top hits again. The README describes this as allowing the agent to page through results and adjust the prompt simultaneously, rather than requiring explicit pagination.
Second, when reading a documentation page, Ref uses the session's accumulated search history to identify which sections of that page are most relevant and drops the less relevant portions before returning. The README gives an example where a complex n8n query results in multiple search and read calls, with each read returning a specific relevant subsection rather than the full page. The goal is to find exactly the context the coding agent needs while using the minimum number of tokens.
Setting Up Ref MCP
There are two deployment modes. The recommended path is the streamable HTTP server hosted by ref.tools, which requires an API key obtained by signing up at ref.tools:
{
"Ref": {
"type": "http",
"url": "https://api.ref.tools/mcp?apiKey=YOUR_API_KEY"
}
}The legacy path runs a local stdio server using npx:
{
"Ref": {
"command": "npx",
"args": ["ref-tools-mcp@latest"],
"env": {
"REF_API_KEY": "<sign up to get an api key>"
}
}
}Both modes require an API key. The current repository contains the stdio server implementation. The streamable HTTP server is hosted infrastructure operated by ref.tools, separate from this repository. For local development, the README documents a dev server and MCP Inspector integration:
npm install
npm run devTwo Tools: ref_search_documentation and ref_read_url
The MCP server exposes two tools to the agent. ref_search_documentation accepts a single required parameter, query, which should be a full sentence or question rather than a keyword string. The README gives example queries like 'Figma API post comment endpoint documentation' and 'n8n merge node vs Code node multiple inputs best practices'. The tool searches both public web documentation and private resources such as repositories and PDFs, depending on what is configured for the account.
ref_read_url accepts a URL and returns that page's content converted to markdown, trimmed to the most relevant sections based on session history. The two tools are designed to be used together: the search tool returns URLs of relevant pages, and the read tool fetches the specific section needed.
For OpenAI deep research compatibility, the server maps these tools to the names search and fetch respectively, since OpenAI's deep research feature expects those specific tool names.
Limitations: Hosted Service Dependency and No Offline Mode
Both deployment modes require an API key and network access to ref.tools. There is no path to run Ref entirely offline or with a purely local index. Teams with documentation hosted in air-gapped environments or who need to avoid sending documentation query metadata to an external service cannot use Ref.
The README does not document which private documentation sources are supported beyond mentioning 'repos and pdfs' as examples. Teams with specific internal documentation formats should verify compatibility before adopting the tool.
context7 is the most directly comparable alternative: it is also an MCP server that provides token-efficient documentation access to AI coding agents. The README includes a comparison in its related search data, suggesting that users evaluating Ref against context7 is a common step. The README does not directly compare the two, but the session-based result deduplication and the page-section trimming based on search history are the features Ref emphasizes as its differentiation from a straightforward fetch-and-return approach.
The package version at the time of the last push was 3.0.3, licensed under MIT. The package.json lists three runtime dependencies: @modelcontextprotocol/sdk at 1.0.3, @modelcontextprotocol/inspector at 0.16.1, and axios at 1.8.4. The inspector dependency is included at runtime, which means the installed bundle is larger than a minimal MCP server would be. Teams running the stdio server as a frequently spawned process may want to verify startup time is acceptable for their agent framework.
Token Cost Framing in the README
The README includes a specific cost calculation to illustrate why token minimization matters for agents using paid API access. Using Claude Opus as a background agent, if a documentation lookup pulls in 10,000 tokens when only 4,000 are relevant, the extra 6,000 tokens cost roughly $0.09 per step at API pricing. A prompt that requires 11 agent steps with that overhead accumulates to $1 in unnecessary cost for that single task. The README presents this as a concrete motivation for the filtering approach, not as a benchmark claim.
The README also cites research from Chroma published as of July 2025 showing that models perform worse as context size increases. The session-based filtering in Ref is designed to address both the cost issue and the accuracy degradation from irrelevant context. Ref tracks which URLs have already been returned in a session and excludes them from subsequent searches, which allows the agent to refine its query across multiple rounds without consuming tokens re-reading content it has already seen. Each new search in a session is therefore more informative than the previous one on the same topic.
Development Setup and Build Process
The server is written in TypeScript. The build produces a CommonJS bundle using esbuild, which is appropriate for an MCP stdio server that needs to start quickly. The package.json defines the build, watch, and inspect workflows:
npm run buildFor development with auto-rebuild:
npm run watchThe MCP Inspector can be launched for visual testing:
npm run inspectA Dockerfile is included for containerized deployment as an HTTP server, with the entrypoint exposing port 8080. The Dockerfile uses a two-stage build: a builder stage installs dependencies and compiles the TypeScript, and a runtime stage copies only the built output and production dependencies. A .mcpbignore file at the root controls which files are excluded from the published package, and a smithery.yaml provides configuration for listing the server on Smithery. The manifest.json describes the server's capabilities for the Smithery registry. The last push to the repository was on 2026-09-14.
Editorial conclusion
Ref MCP is the right choice for engineering teams using Claude Code, Cursor, or other MCP-compatible coding agents who are spending significant token budget on documentation lookups that pull in large, mostly irrelevant pages. The streamable HTTP setup requires signing up for an API key at ref.tools; the local stdio server in this repository is the legacy path. Before choosing between Ref and context7, which serves a similar purpose, check whether Ref indexes the specific private or internal documentation your team uses, since that determines whether the session-based filtering and page truncation features are worthwhile.
Frequently asked questions
Does Ref MCP require an API key?
Yes. Both the streamable HTTP and stdio deployment modes require an API key from ref.tools. The README shows the key as a required field in both configuration examples.
What is the difference between Ref MCP and context7?
Both are MCP servers that provide documentation access to AI coding agents with a focus on reducing context window usage. The README notes Ref's session-based filtering and private documentation support as its design focus, but does not include a direct feature comparison with context7.
What documentation sources can Ref MCP search?
The README states that ref_search_documentation can search public documentation on the web and GitHub, as well as private resources such as repositories and PDFs configured for the account.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/ref-tools-ref-tools-mcp)
Community notes