# OpenBMB/UltraRAG: an MCP-based low-code framework for RAG pipelines

> UltraRAG turns retrievers, generators and evaluation steps into independent MCP servers that a YAML workflow orchestrates. It suits teams that want to iterate on pipeline logic without writing glue code, and it assumes you are comfortable with Python packaging and a CUDA-capable host.

**OpenBMB/UltraRAG** — A Low-Code MCP Framework for Building Complex and Innovative RAG Pipelines

- Repository: https://github.com/OpenBMB/UltraRAG
- Website: https://ultrarag.github.io/
- Stars: 5,698 · Forks: 448
- Language: Python
- License: Apache-2.0
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/openbmb-ultrarag

## The problem UltraRAG targets: glue code between RAG components

A RAG pipeline is usually a retriever, a prompt template, a generator and a scoring step wired together in Python. Every experiment changes one of those parts, and every change means editing the orchestration code that surrounds them. UltraRAG's answer is to make each part a separate MCP server and to move the wiring into YAML. The README describes the project as the first lightweight RAG development framework based on the Model Context Protocol architecture, launched by THUNLP at Tsinghua University, NEUIR at Northeastern University, OpenBMB and AI9stars. It is aimed at two groups: researchers who need to compare pipeline variants, and teams building an industrial prototype who want a demo before committing to an architecture. The pitch is less code, not more abstraction for its own sake. The trade-off is real: you gain configuration-driven control flow, and you take on an MCP runtime and a set of server processes as part of your stack.

## How the MCP server and client split actually works

The architecture is a client-server split. Core RAG components such as Retriever and Generation are standardized as independent MCP Servers. The MCP Client holds the workflow and orchestrates those servers. Control structures that normally require Python, meaning sequences, loops and conditional branches, are expressed in YAML configuration files instead. That is the whole mechanism: the client decides what to call and when, and each server exposes function-level Tools that the client invokes. The README states that new features only need to be registered as function-level Tools to integrate into workflows. The repository layout matches this description. There is a servers/ directory alongside src/, and the pyproject.toml splits optional dependencies into retriever, generation, evaluation, corpus and litellm groups, which means a deployment can install only the servers it runs. A separate ui/ directory holds the frontend, which the Dockerfile builds from ui/frontend with npm before the Python package is installed. The practical consequence of this design is that a pipeline change is often a YAML edit rather than a code change, and the cost is that debugging spans two processes instead of one.

## Installing UltraRAG and running the UI on port 5050

The package requires Python 3.11 or 3.12, since pyproject.toml sets requires-python to >=3.11, <3.13. The project uses uv for dependency resolution; uv.lock is checked in at the repository root, and the Dockerfile copies the uv binary from the astral-sh/uv image before syncing. If you want the full dependency set including the generation and retriever extras, the Dockerfile shows the command the maintainers use:

```bash
uv sync --frozen --no-dev --all-extras
```

The --frozen flag installs exactly what uv.lock pins, and --all-extras pulls in every optional group. On a CPU-only machine this is the wrong command, because the generation extra pins vllm==0.22.1 and the retriever extra pulls faiss-gpu-cu12 on non-Darwin platforms. The pyproject.toml comments explain that LiteLLM is kept separate from the heavy generation extra so that CPU and proxy users can install only the API client. Once dependencies are in place, the container's entry point shows how the web UI is started:

```bash
ultrarag show ui --port 5050 --host 0.0.0.0
```

The Dockerfile exposes port 5050 and uses exactly this command as its CMD, so a container started from the shipped image serves the UI on that port. The README describes the UI as a visual RAG IDE with a Pipeline Builder that supports bidirectional synchronization between canvas construction and code editing, plus a knowledge base component for document Q&A. What you should see after the command is the UltraRAG UI, not a chat window.

## Where UltraRAG is the wrong tool

The dependency file is the clearest statement of the project's assumptions. The generation extra pins vllm==0.22.1 and xgrammar<0.2.4, with comments noting that newer cu129 wheels pull CUDA 13 tooling or conflict with MinerU, and that 0.2.4 has no CPython 3.12 Linux x86_64 wheel. That is a narrow, carefully held position, and it means a GPU host with a different CUDA stack may not install cleanly. The retriever extra pulls faiss-gpu-cu12 on non-Darwin platforms and faiss-cpu only on Darwin, so macOS users get a different retrieval backend by default. Python 3.13 is excluded outright. If your team runs a managed vector database with its own SDK and no interest in MCP, the server abstraction adds a process boundary without removing work. If you need a pipeline that is stable for years with a documented upgrade path, note that the release cadence visible here is three releases in 2026, v0.3.0 in January, v0.3.0.1 in March and v0.3.0.2 in April, and the README does not document a rollback procedure or a deprecation policy for the YAML schema. The last push to the repository was on 2026-09-09, so the codebase is moving, but movement is not the same as a support commitment.

## How UltraRAG differs from FlashRAG, RAGFlow and LightRAG

The comparison that matters is where the orchestration logic lives. FlashRAG is a research toolkit built around Python pipeline classes, so a new control structure means subclassing or editing pipeline code. UltraRAG moves that structure into YAML and delegates execution to MCP servers, which is why the README frames the project around conditional branches and loops expressed in configuration. RAGFlow and LightRAG sit at the other end: they are application-level systems with their own document handling and retrieval stack, and you adopt their pipeline rather than composing your own from parts. UltraRAG does not ship a finished application; it ships the parts and the wiring, plus an evaluation layer. The pyproject.toml evaluation extra contains rouge-score and pytrec-eval-terrier, and the README claims built-in standardized evaluation workflows with ready-to-use benchmarks. Where FlashRAG gives you a Python library to import, UltraRAG gives you servers to run and a YAML file to edit. That is a genuine difference in approach, not a marketing one, and it decides which of the three fits your team.

## Licence, extras and what an upgrade costs

UltraRAG is licensed under Apache-2.0, with the licence text at LICENSE.txt. That permits commercial use and modification, and it includes a patent grant. It does not remove the obligations of the dependencies you install alongside it, and the extras pull in a wide range: vllm, torch, transformers, faiss-gpu-cu12, pymilvus, qdrant-client, mineru[core] and litellm, among others. Each of those carries its own licence, and several are heavier than UltraRAG itself. The upgrade cost is dominated by the pins. Because uv.lock is committed and the Dockerfile uses --frozen, an upgrade means regenerating the lock and re-resolving constraints that the pyproject.toml comments describe as deliberately tight, including the fastmcp range of >=3.3.1,<3.4.5 and numpy>=2,<2.4. A team that installs only the litellm extra avoids vllm and torch entirely and has a much smaller surface to upgrade. This is not legal advice; check the dependency licences against your own policy.

## Conclusion

Adopt UltraRAG if you are prototyping RAG research or an internal demo and want retriever, generation and evaluation steps decoupled into MCP servers you can recombine from YAML. Do not adopt it if you need a supported long-term service with a documented upgrade path, or if you cannot run Python 3.11 or 3.12 with the optional CUDA extras. Before committing, verify which extras your hardware actually needs, since pyproject.toml separates retriever, generation, corpus and litellm, and check the ui/frontend build path in the Dockerfile, because the container builds the frontend with npm before installing the Python package.

## FAQ

### What is RAG and how does it work?

RAG combines retrieval with generation: a retriever finds relevant passages from a knowledge base, and a language model generates an answer conditioned on them. UltraRAG standardizes those parts as independent MCP servers so the retrieval and generation steps can be swapped and recombined from a YAML workflow.

### What is the difference between RAG and MCP?

RAG is a technique for grounding generation in retrieved documents. MCP is the Model Context Protocol, an architecture for exposing tools to a client. UltraRAG uses MCP as the transport and composition layer for RAG components, which is why the README describes it as a RAG development framework based on MCP.

### What does RAG stand for in operating systems?

That is a different meaning of the acronym and it is not what UltraRAG addresses. UltraRAG is a Python framework for retrieval-augmented generation pipelines, and its documentation uses RAG only in that sense.

### Which component is not part of a typical RAG pipeline?

UltraRAG's README names Retriever and Generation as the core components it standardizes as MCP servers, and the pyproject.toml adds an evaluation extra for scoring. Anything outside retrieval, generation and evaluation is not part of the pipeline the project describes.

## Sources

- [License: Apache-2.0](https://github.com/OpenBMB/UltraRAG/blob/main/LICENSE)
- [OpenBMB/UltraRAG on GitHub](https://github.com/OpenBMB/UltraRAG)
- [Project website](https://ultrarag.github.io/)
- [README](https://github.com/OpenBMB/UltraRAG/blob/main/README.md)
- [Releases](https://github.com/OpenBMB/UltraRAG/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/openbmb-ultrarag
