CodeGraphContext: an MCP server that turns a repository into a queryable graph
An MCP server plus a CLI tool that indexes local code into a graph database to provide context to AI assistants.
At a glance
- What is it?
- CodeGraphContext indexes local code with tree-sitter and SCIP into a graph database, then exposes it to AI assistants over MCP or to humans through a CLI. Here is how the pieces fit, how to install it, and where it stops being the right tool.
- Who is it for?
- Adopt CodeGraphContext if you already run an MCP-capable editor and you keep asking structural questions that grep answers badly: call chains, class hierarchies, module boundaries. Skip it if you only need exact string lookup, or if you cannot accept a full indexing pass before the first useful answer, or if your team will not run a graph database alongside the editor.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 3 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem CodeGraphContext solves, and who has it
Text search answers one question well: where does this string appear. It answers badly the questions maintainers actually ask. Which functions reach this one? What implements this interface? Which modules import that package, and through what path? Those answers live in relationships, not in lines, and reconstructing them by hand means opening files, following imports, and holding a mental model that goes stale the moment you switch branches.
CodeGraphContext targets that gap. The README frames it as being built for the moment when plain text search stops being enough, and the comparison table in the README puts it against grep and against RAG over code chunks: string search misses relationships, chunk retrieval can lose symbol-level precision, and a graph trades that for an indexing step. The audience is narrow but real. It is for engineers working in a repository large enough that call chains cross files, and for teams whose AI assistant keeps giving shallow answers because it only sees the files that happen to be open. The MCP server exists specifically so an assistant can query the same structural context a maintainer would build in their head.
How the index is built: tree-sitter, SCIP, and a graph database
The pipeline in the README's flow diagram runs source code into tree-sitter and SCIP indexers, then into a graph database, and out through two consumers: the CLI for direct queries and the MCP server for AI assistants.
Tree-sitter does the parsing. The dependency list pins tree-sitter between 0.24.0 and 0.26.0 and pulls tree-sitter-language-pack, with a separate tree-sitter-c-sharp package. That is the mechanism behind the README's claim that the CLI parses tree-sitter nodes to build the graph: the graph is derived from syntax trees, not from string matching. SCIP indexing is optional and is documented as such, which matters because it is the path to more precise symbol resolution when you need it.
The graph itself is not fixed to one store. The .env.example lists the backend selection values as falkordb, falkordb-remote, neo4j, kuzudb and ladybugdb, with falkordb as the default in the Docker Compose environment. Embedded backends mean you can run without a separate server; the compose file also ships a falkordb service and a neo4j service behind profiles. The pyproject.toml comments are unusually candid about one cost of that choice: FalkorDB and its embedded sibling falkordblite co-evolve with redis-py, and the comment records that redis-py 6 and later breaks the Unix-socket path FalkorDB Lite uses because UnixDomainSocketConnection lacks a host attribute. That is a real pin, not a preference, and it is the kind of detail that decides whether an upgrade is safe.
Installing CodeGraphContext with pip and running a first index
The README's own framing is install with pip and unlock the CLI. The package name on PyPI is codegraphcontext, and the project requires Python 3.10 or newer according to pyproject.toml.
pip install codegraphcontextAfter that, the CLI is the entry point for indexing and querying. The README shows indexing as a CLI operation that parses tree-sitter nodes into the graph.
cgc index .The exact subcommand surface is not reproduced in the README excerpt, so check `cgc --help` after install rather than assuming flags. What you should expect from a successful run is a populated graph for the directory you pointed at, and then the same graph answering structural queries through the CLI.
For a container-based setup, the repository ships a Dockerfile and a docker-compose.yml. The compose file mounts your working directory at /workspace, persists state in a cgc-data volume at /home/cgc/.codegraphcontext, and defaults DEFAULT_DATABASE to falkordb. The Neo4j profile refuses to start without a password.
cp .env.example .env
# set NEO4J_PASSWORD if you use the neo4j profile
docker compose --profile falkordb up -dOne detail worth copying rather than inventing: the FalkorDB service publishes 127.0.0.1:6379:6379, and Neo4j publishes 127.0.0.1:7474:7474 and 127.0.0.1:7687:7687. They are bound to loopback on purpose. The .env.example states the policy directly: bind database ports to 127.0.0.1 in templates and expose via a reverse proxy in production.
Connecting the MCP server to an assistant
The second consumer of the graph is the MCP server, which is what lets an AI assistant query call chains in natural language instead of reading files one at a time. The README labels the project MCP compatible and lists AI-powered code understanding as the reason the server exists.
Configuration is client-side: you register the server with whatever MCP-capable editor or client you use, and the client then sees graph queries as tools. The repository has no committed client configuration file to copy from, so the honest instruction is to follow the project's own MCP documentation and the client's server-registration format rather than a snippet invented here. What you can verify from the repository is that the server side is exercised: there is a test_real_sse.py at the top level, and the GitHub Actions badges cover both a test workflow and an e2e-tests workflow.
Expect the assistant's answers to change character rather than become uniformly better. Graph queries are good at reachability and containment. They are not good at intent, naming conventions, or why a function exists. The README's own comparison table implicitly concedes this: it positions the graph as the tool for repository-wide reasoning, not as a replacement for reading code.
Where CodeGraphContext is the wrong tool
The indexing step is the first limitation, and the README names it as the tradeoff in its own comparison table. Anything that must work on a fresh clone in under a second, or inside a pre-commit hook, is a poor fit. The graph has to exist before it can answer.
The second limitation is the backend matrix. Five backend values appear in .env.example, and the pyproject.toml comment about redis-py shows that at least one of them has a dependency relationship tight enough that a routine transitive upgrade can break the embedded Unix-socket path. If your environment pins redis-py independently, that interaction is yours to manage. The embedded buffer pool is also capped: .env.example documents CGC_EMBEDDED_BUFFER_POOL_MB with a default of 4096 MiB and notes that 0 falls back to roughly 80 percent of system memory. On a large monorepo, that cap is a real ceiling unless you change it.
The third limitation is maturity signalling. pyproject.toml classifies the project as Development Status 3 - Alpha, while the README's Project Details block still reports Version 0.5.1 against a pyproject version of 0.6.13 and a most recent release tag of v0.5.7. Those three numbers do not agree, and the README does not document a rollback path if an index goes wrong. None of that makes the project unusable, but it means you should pin a version and treat upgrades as a deliberate act. If what you actually need is exact string lookup across a small repo, grep is faster and has no state to corrupt.
How it differs from chunk-based retrieval
The obvious alternative is retrieval over embedded code chunks, the pattern most editor assistants use by default. The difference is what gets stored. Chunk retrieval stores text and a vector, so a query returns passages that are semantically near your question. CodeGraphContext stores nodes and edges, so a query returns a path: this function calls that one, this class extends that one, this module imports that package.
That distinction decides which failures you get. Chunk retrieval degrades gracefully: even a mediocre match gives the model something to read, and you can always widen the search. A graph answers precisely or not at all; if the edge was not extracted, the query returns nothing and there is no partial credit. In exchange, the graph does not hallucinate a call chain from plausible-looking prose, because the chain either exists in the store or it does not.
The practical consequence is that the two are complementary rather than competing. A graph backend is worth the indexing cost when your questions are structural and your repository is large. Chunk retrieval remains the better default when your questions are semantic, when the codebase is small enough to fit in context, or when you cannot afford the setup step at all.
Licence, maintenance and the cost of upgrading
The licence is MIT, declared in both the README and the pyproject.toml classifier. That is permissive and imposes no copyleft obligation on your own code. It also means no warranty, and the README points to the LICENSE file for the terms rather than restating them. Nothing here is legal advice; if you redistribute the project inside a product, read the LICENSE yourself.
Maintenance is active by the only measure available: the repository is not archived, and the last push was on 2026-09-06. Releases have been frequent, with v0.5.5 on 2026-08-02, v0.5.6 on 2026-08-05 and v0.5.7 on 2026-08-08. The README lists a single maintainer, Shashank Shekhar Singh, which is worth weighing: a project with one primary maintainer and a fast release cadence is responsive today and fragile if that person steps away.
Upgrade cost is concentrated in the backend drivers rather than the application code. The pyproject.toml comment ties FalkorDB and falkordblite to specific redis-py ranges, and the embedded backends carry a buffer pool setting that may need retuning as your repository grows. There is also a version-reporting inconsistency to resolve before you pin anything: the README says 0.5.1, pyproject.toml says 0.6.13, and the newest release tag is v0.5.7. Check the installed package version with `pip show codegraphcontext` rather than trusting the README block, and read the release notes for the version you actually install.
Editorial conclusion
Adopt CodeGraphContext if you already run an MCP-capable editor and you keep asking structural questions that grep answers badly: call chains, class hierarchies, module boundaries. Skip it if you only need exact string lookup, or if you cannot accept a full indexing pass before the first useful answer, or if your team will not run a graph database alongside the editor. Before committing, verify three things: that the embedded default backend is the one you want selected via DEFAULT_DATABASE, that your target languages are covered by the tree-sitter language pack the project depends on, and that the release you pin actually matches the version reported in pyproject.toml, since the README's Project Details block and the packaging metadata disagree today.
Frequently asked questions
Is CodeGraphContext an MCP server?
It is both an MCP server and a CLI toolkit. The README describes indexing local code into a graph database that serves AI assistants through MCP and developers through the CLI.
What is code graph context?
The project turns a repository into a graph of files, symbols, calls, inheritance and imports, so questions about how code connects can be answered from the graph instead of from text search.
How do I install CodeGraphContext?
The README's quick start installs it from PyPI with pip install codegraphcontext. Python 3.10 or newer is required according to pyproject.toml, and a Dockerfile plus docker-compose.yml are also provided.
Which graph database does CodeGraphContext use?
.env.example lists the backend selection values as falkordb, falkordb-remote, neo4j, kuzudb and ladybugdb, with falkordb as the default in the Docker Compose environment.
Does CodeGraphContext need a remote service?
The README's FAQ states that it works locally with embedded backends, or with external graph databases when you want them. Docker Compose ships both a FalkorDB service and a Neo4j service behind profiles.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/codegraphcontext-codegraphcontext)