Self-hosted service
shinpr/mcp-local-rag avatar
shinpr/mcp-local-rag

shinpr/mcp-local-rag: a local hybrid RAG server for MCP clients and the terminal

Local-first RAG server for developers. Semantic + keyword search for code and technical docs. Works with MCP or CLI. Fully private, zero setup.

405 stars75 forksTypeScriptMIT

At a glance

What is it?
It indexes PDF, DOCX, Markdown and text files on your own machine, then serves semantic plus keyword search over MCP or the CLI. The design is genuinely private, but the supported file set is narrower than the feature list suggests.
Who is it for?
Adopt mcp-local-rag if your documents are PDF, DOCX, Markdown or plain text and policy forbids sending them to a hosted embedding API, and you already run an MCP client such as Claude Code, Cursor, Codex or OpenCode. Do not adopt it if you need to search source code by file extension, spreadsheets, slides or standalone images; the README lists those as unsupported by file ingestion, and HTML must be fetched by the client and passed to ingest_data.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem mcp-local-rag solves, and who it is for

Hosted embedding APIs are the usual way to build retrieval over a document set, and they are the wrong answer for a lot of teams. The README states the motivation plainly: some document sets cannot be sent to a hosted embedding service because of confidentiality or organizational policy, and keeping the index local makes them searchable without adding a per-query API cost.

The audience follows from that. This is a tool for a developer who already has an MCP-capable coding client and a folder of technical documents, and who wants those documents to be searchable inside that client. The README frames it as "Local-first RAG server for developers" and lists Claude Code, Codex, OpenCode and Cursor as example hosts. A second audience is the terminal user: the same index is reachable through a CLI, so you can ingest and query without any MCP client at all.

The second problem is retrieval quality on technical text. Semantic search alone tends to miss exact identifiers. The README makes this argument directly: "Semantic search alone can miss exact identifiers that matter in technical documentation. Keyword reranking keeps those terms visible without giving up natural-language retrieval." That is the reason the project ships hybrid search rather than pure vector search, and it is the most defensible design decision in the repository.

How the indexing and search pipeline actually runs

Everything the README describes happens in one Node.js process. The server parses documents, computes embeddings, stores vectors and runs search locally. After the initial model download, the README says text ingestion and search work offline. There is no external database and no Docker container in the picture; the package.json dependency list points at LanceDB and Transformers-style tooling, which is consistent with an embedded vector store plus an in-process embedding model.

Chunking is the part worth paying attention to. The README calls it semantic chunking: documents are split at topic boundaries rather than fixed character counts, and Markdown code blocks stay intact. For a folder of API references, that matters more than the embedding model choice, because a chunk that cuts a code sample in half produces a useless retrieval hit no matter how good the vectors are.

Search combines two signals. Semantic retrieval finds related concepts; keyword matching boosts exact technical terms such as API names, class names and error codes. The embedding model is configurable, and the README tells you to choose a Hugging Face model that fits the language and domain of your documents. That is a real degree of freedom, and also a real decision you have to make rather than inherit.

Ingestion is exposed through a set of MCP tools. sync_start reconciles the index with configured roots, sync_status polls a running job, ingest_file handles a single file, ingest_data takes text, Markdown or HTML the client already holds, query_documents runs the hybrid search, read_chunk_neighbors pulls surrounding chunks from a result, list_files shows supported files and their ingestion state, delete_file removes an entry, and status reports index and search state. The sync model is incremental: sync_start ingests new and changed files, skips byte-identical files, and removes index entries for files that no longer exist.

Installing mcp-local-rag and running a first query

The requirements are short: Node.js 22 or later, internet access on first use to download the npm package and the embedding model, and a directory containing the documents you want to search. You set BASE_DIR to that directory, and the README notes it is also the security boundary for file operations. Use an absolute path.

For Claude Code, the README gives a single registration command:

bash
claude mcp add local-rag --scope user --env BASE_DIR=/absolute/path/to/your/documents -- npx -y mcp-local-rag

Other hosts take the same server under their own configuration format. For Codex, the README shows a block in ~/.codex/config.toml:

toml
[mcp_servers.local-rag]
command = "npx"
args = ["-y", "mcp-local-rag"]

[mcp_servers.local-rag.env]
BASE_DIR = "/absolute/path/to/your/documents"

For Cursor, the equivalent goes in ~/.cursor/mcp.json:

json
{
  "mcpServers": {
    "local-rag": {
      "command": "npx",
      "args": ["-y", "mcp-local-rag"],
      "env": {
        "BASE_DIR": "/absolute/path/to/your/documents"
      }
    }
  }
}

After restarting the client, you build the index by asking for it in natural language. The README's example prompt is: Sync all documents in the configured root and wait until it finishes. The first sync downloads the default embedding model, about 90 MB, and may take 1 to 2 minutes before ingestion starts. Later runs use the local cache. Once it completes, you search the same way: What does the API documentation say about authentication?

If you would rather skip MCP entirely, the CLI takes two commands:

bash
npx mcp-local-rag ingest ./docs/
npx mcp-local-rag query "authentication API"

The CLI uses the current directory as its document root by default, so run both commands from the same directory or set BASE_DIR and DB_PATH explicitly. That default is convenient and also the easiest way to end up querying an index you did not mean to build.

What mcp-local-rag will not index

The supported-content table is the most important page in the README, and it is narrower than the phrase "search private documents" implies. File ingestion covers PDF, DOCX, TXT and Markdown. HTML is supported only through ingest_data, and only when an MCP client has already fetched the page; the README states that HTML fetching is not built into the server. Plain text or Markdown held in memory also goes through ingest_data with a stable source identifier.

Excel, PowerPoint, standalone images and source-code file extensions are explicitly not supported by file ingestion. For a repository named around developer documentation, the source-code exclusion deserves a second look. If your plan was to point BASE_DIR at a monorepo and search the implementation, that plan does not work here. You would have to convert the code to a supported format first, which changes what the embeddings see.

PDF has a partial escape hatch. PDFs can optionally use a local vision model to describe figures, but the README is careful to say this is not OCR and not image search. So a scanned PDF is still a scanned PDF. The feature describes figures in documents that already have extractable text.

There is a second, quieter limitation around sync state. Only one sync job is retained by the server process: a newer job replaces a finished record, and restarting the server discards it. If you want a durable record of what was indexed and when, you are reading it out of list_files and status, not out of a job history. And a changed PDF keeps the visual profile it was indexed with; sync_start cannot change it. STORE_IMAGES=true, set in the MCP server environment, stores supported PDF and DOCX images for new or changed files selected by sync, but unchanged files remain skipped.

mcp-local-rag versus a hosted RAG pipeline

The obvious alternative is a hosted retrieval stack: send chunks to a hosted embedding API, store the vectors in a managed vector database, and call that from your application. The difference is not just where the bytes live. A hosted pipeline typically gives you larger embedding models, managed scaling, and no local model download, at the cost of a per-query bill and a data-processing question you have to answer before you start. mcp-local-rag inverts that trade: the README's stated reason for existing is that some document sets cannot leave the machine, and the price is that you run the embedding model yourself and accept its quality and speed on your own hardware.

The second alternative is a general-purpose RAG framework that you wire into your own service. That path gives you control over chunking, retrieval and reranking, and it also gives you the work of building all of it. mcp-local-rag's semantic chunking and hybrid retrieval are opinionated defaults you do not get to tune from the outside; the README documents a configurable embedding model, not a configurable chunker or a configurable keyword weight.

The third alternative is simply not using retrieval at all and pasting documents into the model's context. For a handful of short files that is faster than any index. It stops working when the corpus outgrows the context window, which is the point at which a chunked, searchable index starts paying for itself.

Maintenance, upgrades and the MIT licence

The repository is not archived, and the last push was on 2026-09-07, which is recent enough that the project is being worked on. The release list supports that: v0.18.2 on 2026-09-05, v0.18.3 on 2026-09-07, and v0.18.4 later the same day. Three releases in three days is a fast cadence, and it cuts both ways. You get fixes quickly; you also get a version number that moves under you, and the 0.x major version signals that the interface is not frozen.

Upgrade cost is mostly the npm package and the embedding model cache. Because the server is launched through npx -y mcp-local-rag, an MCP host will pull whatever the latest published version is on the next start unless you pin it. That is worth knowing before you put it in a team configuration file. The README does not document a rollback procedure or a version-pinning flag for the MCP host configuration, so pinning is something you would have to arrange in your own command line.

The licence is MIT, which is permissive and short. That is a statement about the repository's LICENSE file, not legal advice; if your organization has rules about which licences may be used in internal tooling, check them against the actual LICENSE text rather than this paragraph. One practical note: the README says the first sync downloads an embedding model from Hugging Face. Whatever licence that model carries is separate from the MIT licence on the server, and the README does not enumerate model licences.

Editorial conclusion

Adopt mcp-local-rag if your documents are PDF, DOCX, Markdown or plain text and policy forbids sending them to a hosted embedding API, and you already run an MCP client such as Claude Code, Cursor, Codex or OpenCode. Do not adopt it if you need to search source code by file extension, spreadsheets, slides or standalone images; the README lists those as unsupported by file ingestion, and HTML must be fetched by the client and passed to ingest_data. Before committing, verify three things on your own corpus: that Node.js 22 or later is available, that BASE_DIR points at the directory you actually want indexed, and that the first sync finishes, since it downloads the default embedding model (about 90 MB) and may take 1 to 2 minutes before ingestion starts.

Frequently asked questions

Is RAG obsolete with MCP?

No. MCP is the transport that lets an AI client call tools such as query_documents; RAG is what makes a document set searchable in the first place. mcp-local-rag uses MCP to expose an index it builds and stores locally, so the two are complementary rather than substitutes.

Is MCP needed for RAG?

Not for this project. The README documents a CLI path with npx mcp-local-rag ingest and npx mcp-local-rag query that uses the same index without an MCP client. MCP is the integration route for AI coding tools, not a requirement for building or querying the index.

How do I use RAG locally with mcp-local-rag?

Set BASE_DIR to the absolute path of your document directory, register the server with your MCP host, restart the client, then ask it to sync the configured root and wait until it finishes. The first sync downloads the default embedding model, about 90 MB, and may take 1 to 2 minutes before ingestion starts.

What does local MCP mean for mcp-local-rag?

The README describes a standard MCP protocol server running over local stdio, so the client talks to a process on your machine rather than a remote endpoint. Document parsing, embeddings, storage and search all run locally, and after the initial model download ingestion and search work offline.

Official sources

  1. Issues
  2. License: MIT
  3. README
  4. Releases
  5. shinpr/mcp-local-rag on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/shinpr-mcp-local-rag.svg)](https://hysenlabs.com/projects/shinpr-mcp-local-rag)