# Code-Graph-RAG: querying a monorepo through a Memgraph knowledge graph

> Code-Graph-RAG parses a multi-language repository with Tree-sitter, stores the structure in Memgraph, and answers questions in natural language. It is a strong fit for polyglot monorepos and a poor fit for anyone who does not want Docker and a graph database in the loop.

**vitali87/code-graph-rag** — The ultimate RAG for your monorepo. Query, understand, and edit multi-language codebases with the power of AI and knowledge graphs.

- Repository: https://github.com/vitali87/code-graph-rag
- Website: https://code-graph-rag.com
- Stars: 5,178 · Forks: 684
- Language: Python
- License: MIT
- Published: 2026-08-04 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/vitali87-code-graph-rag

## The monorepo problem Code-Graph-RAG targets

Embedding-based retrieval treats a codebase as text. It splits files into chunks, stores vectors, and returns the chunks that look closest to your question. That works when the answer lives in one file. It fails when the answer is a relationship: which services call this function, what implements this interface, what breaks if this class changes. A vector store has no edge to traverse, so the model reconstructs call chains from whatever chunks happened to rank highest.

Code-Graph-RAG takes the other route. The README describes a Tree-sitter parser that extracts functions, classes, methods, modules and the relationships between them, storing all of it in Memgraph under what the project calls a single language-agnostic schema. The audience is teams with a monorepo in more than one language, where a Python service calls a Go worker and a TypeScript frontend consumes both. The project also exposes an MCP server, so an editor or agent host can query the same graph rather than a separate index.

## Parser, graph, Cypher: the path a question takes

The README gives the data flow directly. Source code goes through the Tree-sitter parser into AST analysis, then into the Memgraph knowledge graph. On the query side, a user question goes to an AI model that generates Cypher, the Cypher runs against the graph, and the results come back as an answer. The RAG system lives in codebase_rag/, described as an interactive CLI.

That split matters for how you debug it. When an answer is wrong, the fault is either the graph (the parser missed an edge, or the file was never indexed) or the generated Cypher (the model wrote a query that does not express the question). You can inspect both, because the intermediate artifact is a database you can open, not a vector space you cannot read. The project also documents an overlay for runtime behaviour: cgr trace records a test run, or pulls production eBPF profiles, and merges the calls that actually happened into the graph. That addresses dispatch that static analysis cannot see, such as dynamic calls resolved only at runtime.

## Installing Code-Graph-RAG and indexing a first repository

The package is published to PyPI and ships two console scripts, code-graph-rag and cgr, both pointing at codebase_rag.cli:app. The README recommends uv and asks for the treesitter-full extra, which pulls in all language grammars, plus semantic for vector search.

```bash
uv tool install "code-graph-rag[treesitter-full,semantic]"
```

pipx works the same way if you prefer it. Before anything runs you need Python 3.12 or newer, Docker for Memgraph, cmake and ripgrep. The README notes the wheel is pure Python, so the package itself installs anywhere with a suitable interpreter, but dependencies may still need platform wheels or build tools, and it names cmake as the requirement for pymgclient. On Raspberry Pi OS Bookworm the piwheels build is listed as failed because that system Python is 3.11, below the floor; the README suggests pinning the interpreter explicitly.

The packaged stack starts without a compose file:

```bash
cgr daemon up
```

Then point it at a repository. The first command parses and indexes, the second opens the query session against the graph that already exists:

```bash
cgr start --repo-path /path/to/repo --update-graph
cgr start --repo-path /path/to/repo
```

Model providers are configured through environment variables. The .env.example lists six combinations, including all-local Ollama, all-OpenAI, Google AI Studio, Vertex AI, and mixed setups where the orchestrator and the Cypher generator use different providers. The keys follow the pattern ORCHESTRATOR_PROVIDER, ORCHESTRATOR_MODEL, ORCHESTRATOR_API_KEY and the matching CYPHER_ prefixed set.

```bash
ORCHESTRATOR_PROVIDER=ollama
ORCHESTRATOR_MODEL=qwen2.5-coder
ORCHESTRATOR_ENDPOINT=http://localhost:11434/v1
```

With that in place, cgr start should drop you into an interactive prompt where a plain-English question is translated into Cypher and answered from the graph. If the graph is empty, the answer will be too, so run the --update-graph pass first and let it finish.

## Version cadence, --clean, and other ways this bites

The versioning scheme is unusual and worth reading before you file a bug. Tags are cut on every merge, but GitHub Releases and PyPI uploads follow every fiftieth version plus any security fix. So the newest tag on main typically runs tens of patch versions ahead of what uv tool install gives you. The README is explicit that nothing is stuck and the cadences differ by design. If you need code newer than the latest release, install from git with the @main reference the README provides. The practical consequence: a bug you see on PyPI may already be fixed on main, and a fix on main may not reach PyPI for a while unless it is a security issue.

The sharper edge is --clean. The README states it deletes every project in the shared graph, not just the one you named, and asks for confirmation first when other projects would be destroyed. The graph is shared across repositories by design, which is what makes cross-project queries possible and also what makes a careless reset expensive. Re-indexing a large monorepo is not free, and there is no documented per-project purge.

Scale is the other unstated boundary. The README does not publish index sizes, indexing times, or memory requirements, so there is no way to estimate from the documentation whether a repository of a given size will fit comfortably. The Dockerfile installs ripgrep and libssl3 in the runtime stage and builds with cmake and build-essential, which tells you the container is doing real compilation work at build time, not just copying a wheel.

## Where a plain vector index is the better choice

The obvious alternative is chunk-and-embed retrieval over the same repository, the approach behind most editor assistants. The difference is structural, not a matter of tuning. A vector index stores text and retrieves by similarity, so it needs no parser, no database, and no schema, and it works on languages Tree-sitter grammars do not cover. Code-Graph-RAG needs a parser that understands your language and a graph that holds the edges; in exchange, questions about call chains and implementations are answered by traversal rather than by guessing which chunks matter.

That trade favours the graph when your questions are relational and your languages are on the supported list. It favours the vector index when your questions are local ("explain this function"), when your code is in a language the project only supports structurally, or when you cannot run a database alongside your editor. The README's own language list draws that line: Python, TypeScript, TSX, JavaScript, Rust, Go, Java, C, C++, C#, PHP, Lua and Dart are fully supported, Scala is in development, and Ruby, Kotlin, Swift, Elixir, Haskell, Solidity, Bash and Nix get structural support through the pluggable ast-grep tier, meaning modules, functions, classes where the language has them, and imports. If your monorepo is mostly Ruby, you are getting imports and function nodes, not the full relationship graph the pitch implies.

## Maintenance, licence and what an upgrade costs

The repository is not archived and the last push was on 2026-08-26, with releases v0.0.770 on the same date, v0.0.720 on 2026-08-22 and v0.0.670 on 2026-08-18. That is a dense release rhythm, and the version cadence described above means the tag history is denser still. The project is classified as Development Status 4 - Beta in pyproject.toml, which is consistent with the 0.0.x version numbers.

Upgrading is a tool install, not a library bump, unless you embed it. Because cgr is a CLI and an MCP server, an upgrade means re-running the install command and restarting the daemon. The graph schema is the risk surface: if a release changes node or relationship properties, an existing graph may need re-indexing, and re-indexing is the expensive operation. The README does not document a migration path between graph versions, so treat a schema change as a full rebuild until you know otherwise.

The licence is MIT, declared in both pyproject.toml and the LICENSE file. MIT is permissive and places few obligations on redistribution, but the article is not legal advice: the model providers you configure, the Memgraph image you run, and the Tree-sitter grammars you bundle each carry their own terms, and those are the ones to check before shipping anything derived from this stack.

## Conclusion

Adopt it if your repository spans several languages and you already run Docker, since the graph is the whole point and a single-language project gains little from it. Skip it if you cannot run Memgraph, if your code must stay off a hosted model endpoint, or if you expect an editor-quality refactor from the agent. Verify first that the packaged stack comes up with cgr daemon up on your machine, that your model provider is reachable from the same environment, and that cgr start --repo-path finishes indexing before you judge any answer.

## FAQ

### What is a graph RAG system?

In this project, it means retrieval that runs against a knowledge graph instead of a vector index: the codebase is parsed into nodes and relationships in Memgraph, and a natural-language question is turned into a Cypher query against that structure. The README describes the pipeline as Source Code to Tree-sitter Parser to AST Analysis to Memgraph, and on the query side User Query to AI Model to Cypher Query to Graph Results to Response.

### Is graph RAG better than vector RAG?

The README does not compare the two, and the answer depends on the question you are asking. Vector retrieval returns text chunks by similarity and needs no parser or database; Code-Graph-RAG returns graph traversals, which suits relational questions such as which functions call a given method, but requires a supported language and a running Memgraph instance.

### How do I install code-graph-rag with Docker?

Docker is a prerequisite for Memgraph, and the README says the packaged Memgraph and Qdrant stack starts with cgr daemon up, with no compose file needed. The CLI itself installs from PyPI with uv tool install or pipx, using the treesitter-full and semantic extras.

### Which languages does code-graph-rag support?

Python, TypeScript, TSX, JavaScript, Rust, Go, Java, C, C++, C#, PHP, Lua and Dart are fully supported. Scala is in development, and Ruby, Kotlin, Swift, Elixir, Haskell, Solidity, Bash and Nix have structural support through the ast-grep tier, covering modules, functions, classes where the language has them, and imports.

### Does code-graph-rag work with local models?

Yes. The .env.example lists an all-Ollama configuration using ORCHESTRATOR_PROVIDER=ollama with ORCHESTRATOR_ENDPOINT=http://localhost:11434/v1, and mixed setups where the orchestrator and the Cypher generator use different providers. OpenAI, Google AI Studio and Vertex AI are the other documented options.

## Sources

- [Official documentation](https://code-graph-rag.com)
- [Official README](https://github.com/vitali87/code-graph-rag#readme)
- [Project repository](https://github.com/vitali87/code-graph-rag)
- [Release notes](https://github.com/vitali87/code-graph-rag/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/vitali87-code-graph-rag
