ApeRAG Review: GraphRAG, MCP and Kubernetes Deployment in One Stack
ApeRAG: Production-ready GraphRAG with multi-modal indexing, AI agents, MCP support, and scalable K8s deployment
At a glance
- What is it?
- ApeCloud's ApeRAG bundles vector, full-text, graph, summary and vision indexes behind a FastAPI service with an MCP endpoint. It is a serious deployment, and the operational weight matches the feature list.
- Who is it for?
- Adopt ApeRAG if you need graph plus vector plus full-text retrieval behind one API and you already run PostgreSQL, Redis, Qdrant, Elasticsearch and Neo4j, or you are willing to let the Helm chart install them. Do not adopt it if you want a single-binary local RAG tool or you cannot accept an alpha version line: the newest releases listed are v0.7.0-alpha.34 from 2026-03-11.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 18 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What ApeRAG solves, and for whom
Most retrieval stacks force a choice. A vector store gives you semantic recall but loses exact identifiers and rare terms. A keyword index handles those but misses paraphrase. A knowledge graph answers questions about relationships that neither index can express. ApeRAG's position is that you should not have to pick: the README lists five index types, Vector, Full-text, Graph, Summary and Vision, and a hybrid retrieval engine that queries across them.
The target reader is a team building an internal question-answering product over a document corpus that includes tables, charts and scanned material. The README names multimodal document processing and vision support for images and charts, and describes an AI agent layer that can pick relevant collections on its own. That combination points at enterprise knowledge bases rather than personal note-taking.
The project is written in Python, licensed Apache-2.0, and its pyproject.toml pins requires-python to >=3.11.12 and <3.13. That upper bound matters: if your platform standardizes on Python 3.13, you are outside the supported range before you start.
Inside the retrieval stack: five indexes and an agent loop
The architecture visible in the repository is a set of cooperating services rather than a monolith. The API service is FastAPI, exposed on port 8000 with documentation at /docs. The frontend is a separate React container on port 3000. Async work runs through Celery, with flower and django-celery-beat present in the dependency list, which implies scheduled and background indexing jobs rather than synchronous document ingest.
Storage is split by index type. Qdrant backs vector search, Elasticsearch backs full-text search, and Neo4j backs the graph. PostgreSQL holds relational data and Redis handles caching and the Celery broker. docker-compose.yml declares named volumes for each of these plus a shared volume mounted at /shared in the API container.
The graph layer is the part worth reading the README on closely. It describes a "deeply modified LightRAG implementation with advanced entity normalization (entity merging) for cleaner knowledge graphs." Entity merging is the unglamorous work that decides whether a graph is usable: the same organization mentioned three ways should collapse into one node. The README claims the modification improves relational understanding, and that is a claim about a fork rather than about LightRAG itself.
On top sits the agent layer. The README says built-in agents use MCP tool support to identify relevant collections, search them, and fall back to web search. That is a routing decision made at query time, and it means retrieval quality depends on the agent choosing the right collection, not only on index quality.
Installing ApeRAG with Docker Compose
The README states minimum requirements of 2 CPU cores, 4 GiB RAM, Docker and Docker Compose. The quick start clones the repository, copies the environment template, and brings the stack up. The --pull always flag means Compose fetches the published images rather than building locally.
git clone https://github.com/apecloud/ApeRAG.git
cd ApeRAG
cp envs/env.template .env
docker-compose up -d --pull alwaysAfter the containers report healthy, the README gives two addresses: the web interface at http://localhost:3000/web/ and the API documentation at http://localhost:8000/docs. The API container's healthcheck retries twelve times at ten second intervals after a thirty second start period, so the README's own configuration expects roughly two minutes before the API is considered up. If you open the browser earlier and see a connection error, that is the expected window, not a failure.
The first real use is connecting an MCP client. The README shows a client configuration pointing at the MCP endpoint with a Bearer token, and states that the header takes priority over the APERAG_API_KEY environment variable as a fallback.
{
"mcpServers": {
"aperag-mcp": {
"url": "http://localhost:8000/mcp/",
"headers": {
"Authorization": "Bearer your-api-key-here"
}
}
}
}Replace the placeholder with a key from your ApeRAG settings, and use your deployed origin instead of localhost if the stack is not running on your machine. The README lists three capabilities the MCP server exposes: collection browsing, hybrid search across vector, full-text and graph methods, and natural language querying.
If you need better handling of tables, formulas and complex layouts, the README describes an optional MinerU-based parsing service, enabled through a Compose profile. The GPU variant is the one the README recommends for that path.
DOCRAY_HOST=http://aperag-docray:8639 docker compose --profile docray up -d
DOCRAY_HOST=http://aperag-docray-gpu:8639 docker compose --profile docray-gpu up -dThe Makefile offers shortcuts, make compose-up WITH_DOCRAY=1 and make compose-up WITH_DOCRAY=1 WITH_GPU=1, but those require GNU Make on the host.
Where ApeRAG gets expensive or awkward
The honest limitation is the dependency surface. A default deployment wants PostgreSQL, Redis, Qdrant, Elasticsearch and Neo4j running before the API will start, and the Compose file encodes that with health-gated depends_on entries. On a laptop that is a lot of memory for a retrieval experiment, and the README's 4 GiB floor describes the minimum for the stack to come up, not the point at which graph indexing over a large corpus is comfortable.
The second limitation is maturity. The releases listed for this repository are all v0.7.0-alpha builds, the most recent being v0.7.0-alpha.34 dated 2026-03-11. The last push to the default branch was 2026-05-02. That is a project still iterating on its own interfaces, and an alpha version line is a poor fit for anyone who needs a frozen API contract across a long integration.
The third is that graph extraction is not free at query time or at ingest time. Building a knowledge graph means calling a language model over your documents, and the README does not publish token costs, throughput figures or a comparison against running vector-only retrieval. If your corpus is large and your budget is fixed, the graph index is the component to model before you commit, not after.
Finally, the README does not document rollback, backup or restore procedures for the individual stores. It tells you how to bring the stack up and how to deploy it to Kubernetes. It does not tell you how to recover a corrupted graph or migrate Qdrant collections between versions. That silence is worth treating as work you own.
ApeRAG compared with plain LightRAG or a vector-only stack
The closest reference point is LightRAG, which ApeRAG's README names directly as the basis of its graph layer. The difference in approach is scope. LightRAG is a retrieval library you embed in your own application and point at a storage backend you choose. ApeRAG wraps a comparable graph approach in a service: a FastAPI backend, a React console, Celery workers, authentication through fastapi-users, audit logging, and an MCP endpoint that external assistants can call.
That means the trade is control for integration. With a library you decide how documents are chunked, when the graph is rebuilt, and which store holds what. With ApeRAG those decisions are largely made for you by the Compose file and the configuration in config/ and envs/. You gain a working console and an MCP server on day one, and you lose the ability to swap Elasticsearch for something lighter without reading the source.
Against a vector-only stack such as a plain Qdrant plus an embedding model, the difference is what questions you can answer. Vector search handles similarity. It does not handle "which suppliers are connected to this component through any number of intermediaries," which is the shape of question a graph index exists to answer. If your queries are all of the first kind, ApeRAG's extra stores are overhead you will pay for in memory and operational attention.
Deployment, licensing and upgrade cost
For production the README recommends Kubernetes over Compose, and points at a Helm chart in the repository's deploy directory. The stated prerequisites are a cluster at v1.20 or later, kubectl configured against it, and Helm v3 or later. The README describes two paths for the databases: install PostgreSQL, Redis, Qdrant and Elasticsearch yourself, or use the provided KubeBlocks integration to have them deployed for you. The first path gives you control over versions and backups. The second is faster and leaves you depending on the chart's choices.
The licence is Apache-2.0, which permits commercial use and modification and includes an explicit patent grant. It also means that if you modify ApeRAG and distribute it, you carry the standard obligations around notices and stating changes. None of that is legal advice, and the LICENSE file in the repository root is the text that governs.
Upgrade cost is where the alpha version line bites. Because the releases are pre-1.0, interface changes can arrive between minor versions, and the Compose file pulls images by tag with a VERSION variable defaulting to v0.0.0-nightly. Pinning an explicit version rather than relying on the default is the difference between a reproducible deployment and one that changes under you at the next pull. The repository ships a uv.lock file, which pins Python dependencies precisely for anyone building from source, but that lock does not constrain the container images you pull.
Editorial conclusion
Adopt ApeRAG if you need graph plus vector plus full-text retrieval behind one API and you already run PostgreSQL, Redis, Qdrant, Elasticsearch and Neo4j, or you are willing to let the Helm chart install them. Do not adopt it if you want a single-binary local RAG tool or you cannot accept an alpha version line: the newest releases listed are v0.7.0-alpha.34 from 2026-03-11. Before committing, check the .env produced from envs/env.template for the LLM and embedding provider keys it expects, confirm that the five index types are all enabled for your collections, and verify that your MCP client can reach http://localhost:8000/mcp/ with a Bearer token.
Frequently asked questions
What are the minimum system requirements to run ApeRAG?
The README states CPU of at least 2 cores, RAM of at least 4 GiB, plus Docker and Docker Compose. Those are the requirements for the Compose quick start, which brings up the API, frontend, Celery worker and the backing databases.
How do I connect an MCP client to ApeRAG?
Point the client at http://localhost:8000/mcp/ and send an Authorization header with a Bearer token from your ApeRAG settings. The README says the HTTP header takes priority, with the APERAG_API_KEY environment variable as a fallback.
Which databases does ApeRAG require?
The README lists PostgreSQL, Redis, Qdrant and Elasticsearch as required, and the dependency list also includes a Neo4j client for the graph index. The docker-compose.yml file defines volumes for all five, so a full local start runs every one of them.
Is ApeRAG stable enough for production?
The releases listed for the repository are v0.7.0-alpha builds, the newest being v0.7.0-alpha.34 from 2026-03-11, and the last push to the default branch was 2026-05-02. The README recommends Kubernetes with Helm charts for production deployment, but an alpha version line means interfaces can still change between releases.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/apecloud-aperag)