Modular RAG MCP Server: a RAG pipeline built to be taken apart
A modular RAG (Retrieval-Augmented Generation) system with MCP Server architecture. Using Skill to make AI follow each step of the spec and complete the code 100% by AI.
At a glance
- What is it?
- A Python reference implementation of hybrid retrieval, reranking, image captioning and evaluation, exposed over MCP, where every stage sits behind an interface and the README admits it is also a portfolio project.
- Who is it for?
- What this repository is good at is being read. The DEV_SPEC file, the three-branch structure and the pluggable interfaces exist so a reader can see how a retrieval pipeline is assembled, and the retrieval design itself, sparse and dense candidates fused with RRF then reranked, is the part worth stealing for a real system.
- Can I use it commercially?
- Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
- Is it still maintained?
- Activity is slowing. The repository last received commits 7 months ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 20, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Two-stage retrieval with an interface at every seam
The retrieval design is the part of this project that holds up under scrutiny, and the README states it plainly: BM25 sparse retrieval handles exact matches on proper nouns, dense embeddings handle semantic matches, the two candidate sets are fused with RRF, and an optional Cross-Encoder or LLM rerank pass refines the ordering. Coarse recall first, precise ordering second. That two-stage arrangement is the standard answer to the problem that a single vector search over technical documents gets wrong in both directions at once.
The modularity claim rests on a specific design decision rather than on adjectives. The README lists LLM, Embedding, Reranker, Splitter, VectorStore and Evaluator as the core seams, each defined as an abstract interface, with the backend swapped through configuration rather than through code edits. If you are evaluating whether this is real, that list is the thing to check in `src/`, because an interface that every implementation also reaches through directly is not a seam.
The pipeline itself is described as PDF to Markdown to Chunk to Transform to Embedding to Upsert, with image captioning folded into the text path: a Vision LLM generates a description for an image and the description is stitched into the chunk. The stated result is that searching text retrieves an image, without a separate multimodal index. `markitdown[pdf]` in the dependency list is what does the PDF conversion, and `chromadb` is the vector store, with `langchain-text-splitters` handling chunking.
Three tools, one dashboard, and six Streamlit pages
The MCP surface is small and specific. The README names three tools: `query_knowledge_hub`, `list_collections` and `get_document_summary`. A server with three well-chosen tools is easier for a model to use correctly than one that exposes forty, so this is a sensible boundary rather than a limited one, and the intent is that clients such as GitHub Copilot and Claude Desktop can call them with no frontend work.
Around that sits a Streamlit dashboard described as six pages: a system overview, data browsing, ingestion management, ingestion tracing, query tracing and an evaluation panel. The two tracing pages are the interesting ones, because they follow the README's framing of observability as white-box tracing where every intermediate state on the ingestion chain and the query chain stays visible. In a retrieval system, the failure you actually need to debug is the one where the document was chunked badly or the reranker dropped the right passage, and tracing is the only way to tell those apart.
Evaluation is the third pillar, and the README is unusually firm about method: Ragas plus custom metrics, supporting golden test set regression, described as a refusal to tune by feel. That last point is the argument worth carrying to any retrieval project, including one you write yourself. A golden test set turns retrieval quality into something that can regress in CI rather than something a reviewer judges by reading five sample answers.
Three branches, three different things to clone
The repository ships three branches and the README is direct about which one you want. `main` is described as the cleanest complete code, always holding exactly one commit, for people who want the finished system quickly or plan to build on top of it. `dev` has identical code with the full commit history, so you can trace how the project was assembled step by step. `clean-start` contains only the skeleton: Agent Skills plus `DEV_SPEC.md`, with all task progress zeroed out.
The README recommends `clean-start` most strongly and suggests going further, deleting DEV_SPEC and designing the document yourself. The stated rationale is that the code is meant to be written by an AI reading the spec, which means the spec is the actual artifact of interest.
That makes this repository an unusual case, because it is explicitly two things at once: a working RAG system and a study in spec-driven development. `DEV_SPEC.md` sits at the repository root as a peer of `README.md`, and the README says to consult it for architecture detail, module descriptions and task scheduling. If you are interested in the retrieval code, `main` is the efficient read. If you are interested in the method, the `dev` branch history is.
One inconsistency to expect: the repository is named MODULAR-RAG-MCP-SERVER, while the quick start instructs you to `cd Modular-RAG-MCP-Server` after cloning, so on a case-sensitive filesystem the directory name you typed will not match what was cloned.
Getting it running through a Skill rather than a shell script
The quick start is two steps. Clone the repository, then let an agent configure the environment:
git clone <repo-url>
cd Modular-RAG-MCP-ServerThe second step is not a command you type into a terminal. The README describes a Setup Skill that handles provider selection, API key configuration, dependency installation, config file generation and Dashboard startup, invoked by opening the project in VS Code and typing into the Copilot or Claude conversation box:
setupThis is worth pausing on, because it means the documented first-run path assumes you have a working MCP-capable agent client and a willingness to hand it API keys. If you would rather drive it directly, the packaging gives you the pieces: `pyproject.toml` declares `mcp-server = "main:main"` as a console script, so the server is reachable as a command once installed.
The dependency list is where the real design shows. Alongside the retrieval stack it includes `jieba`, which is a Chinese word segmenter, alongside `markitdown[pdf]` for document conversion, `datasets` for evaluation data, `ragas` for scoring, `streamlit` for the dashboard and `mcp` for the protocol itself. Python 3.10 is the floor. The dev group adds pytest with asyncio, coverage and mock support, mypy and ruff, plus `openai`, and the pytest configuration defines markers for unit, integration, e2e, llm and slow tests, including an `llm` marker meant to be deselected when you do not want real API calls in your test run. That marker is the detail to notice: the project anticipates a test suite that normally talks to a paid model.
Packaging signals that this is not a service yet
A few fields in `pyproject.toml` set expectations honestly. The version is 0.1.0 and the development status classifier is 3, Alpha, which is at odds with the README's confidence about the architecture. The author entry is still the template placeholder, a name and an example.com address, rather than the person who wrote it. Neither fact is a defect, and both tell you the packaging was scaffolded rather than released.
The license is a similar story. The project file declares MIT, and the classifiers include the OSI approved MIT license line, while GitHub reports no license for the repository and there is no LICENSE file in the tree. The MIT declaration in the package metadata is the stronger of the two signals, but if you plan to redistribute this code rather than read it, the absence of a LICENSE file at the root is the gap to raise with the author first.
The repository has no releases and no topics. The last push was on 2026-03-10, and the description on the repository is explicit about what the project is for: a modular RAG system with MCP architecture, and a practical project plus teaching material designed for people studying for large-model roles.
What that description makes unambiguous is the audience. The README spends a substantial part of its length on how to use this for interview preparation, resume writing and building a proof of concept, including a suggested wording for a resume bullet and a section addressed to product managers. Whether that framing is a liability depends entirely on why you are opening the repository: if you are looking for infrastructure to run, you will find the design useful and the packaging thin; if you are studying retrieval engineering or preparing for a role, the framing is the feature.
What to take from it and what to leave
Compare this to what you would otherwise reach for. A framework such as LangChain gives you retrievers, loaders and chains off the shelf, with a large ecosystem and a learning curve that arrives alongside the abstractions. A hosted retrieval service gives you indexing and search without the operational work, at a recurring cost and with your documents leaving your infrastructure. This repository sits between them: it assembles the stages explicitly so you can see them, using ChromaDB and an MCP server rather than a chain abstraction.
The parts worth copying into real work are the two-stage retrieval with RRF fusion, the interface-per-stage structure, the golden test set discipline, and the separation of ingestion and query tracing. Those are durable decisions, and the DEV_SPEC file is the place to read how they were specified.
The parts to leave are the Skill-driven setup and the alpha packaging. A production deployment wants pinned versions, a real author and license file, and credentials handled by your own secret store rather than typed into an agent conversation. There is also a dependency to weigh before building on this: `jieba` in the required list indicates the retrieval path is written with Chinese tokenisation in mind, which is a deliberate choice if your corpus is Chinese and a poor default if it is not, since BM25 quality depends on segmentation.
Editorial conclusion
What this repository is good at is being read. The DEV_SPEC file, the three-branch structure and the pluggable interfaces exist so a reader can see how a retrieval pipeline is assembled, and the retrieval design itself, sparse and dense candidates fused with RRF then reranked, is the part worth stealing for a real system. What it is not is a deployed service: the package is version 0.1.0 and marked as alpha, the author field in `pyproject.toml` is still a placeholder, and the last push was on 2026-03-10. Copy the module boundaries and the two-stage retrieval shape into a codebase that already has embeddings configured, rather than wiring your own product directly to this one.
Frequently asked questions
What is an MCP server vs RAG?
They solve different problems and this repository combines both. RAG is the retrieval technique, searching a document collection and putting the results into a prompt, while an MCP server is a protocol that exposes tools to an AI client. Here the RAG pipeline is the implementation and MCP is the interface, with three tools exposed: query_knowledge_hub, list_collections and get_document_summary.
What is meant by an MCP server?
A server that speaks the Model Context Protocol, publishing tools that an MCP client such as GitHub Copilot or Claude Desktop can call. In this project it wraps the retrieval pipeline so an assistant can query a knowledge base directly, with no separate frontend to build.
Does Modular RAG MCP Server support hybrid search?
Yes. The README describes BM25 sparse retrieval for exact matches on proper nouns combined with dense embedding retrieval for semantic matches, fused with RRF, then optionally reranked with a Cross-Encoder or an LLM to trade recall against precision.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/jerry-ai-dev-modular-rag-mcp-server)