# codebase-memory-mcp: a single-binary code intelligence server for AI coding agents

> codebase-memory-mcp indexes a repository into a persistent knowledge graph and exposes it to coding agents over MCP. The trade-off is real: a native binary that reads your codebase and rewrites your agent config, with 15 tools and a local 3D graph UI.

**DeusData/codebase-memory-mcp** — High-performance code intelligence MCP server. Indexes codebases into a persistent knowledge graph, average repo in milliseconds. 158 languages, sub-ms queries, 99% fewer tokens. Single static binary, zero dependencies.

- Repository: https://github.com/DeusData/codebase-memory-mcp
- Website: https://deusdata.github.io/codebase-memory-mcp/
- Stars: 45,470 · Forks: 3,722
- Language: C
- License: MIT
- Published: 2026-08-08 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/deusdata-codebase-memory-mcp

## The problem codebase-memory-mcp targets

A coding agent exploring an unfamiliar repository does the same thing a new engineer does without a map: it greps, opens files, reads, and repeats. The README frames the cost in tokens, quoting about 412,000 tokens for file-by-file search against roughly 3,400 tokens for five structural queries. Whether or not that ratio holds for your repository, the underlying complaint is familiar. File reads are stateless. The agent relearns the call graph on every session because nothing persisted it.

codebase-memory-mcp answers that by building a graph once and querying it many times. The intended user is someone running an agent against a codebase large enough that repeated exploration is visible in cost and latency, and who wants structural answers (who calls this, what breaks if I change it, which HTTP routes exist) rather than text matches. The README also positions it for infrastructure-as-code work: Dockerfiles, Kubernetes manifests and Kustomize overlays are indexed as graph nodes, with Resource nodes for K8s kinds and Module nodes for Kustomize overlays linked by IMPORTS edges.

It is not a linter, a type checker or a build tool. It does not compile your project. It reads it.

## How the index is built and what the agent sees

The pipeline described in the README is RAM-first. Parsing runs through tree-sitter AST analysis across the supported languages, with vendored grammars compiled into the binary rather than loaded from disk. On top of that, a Hybrid LSP layer adds semantic type resolution for Python, TypeScript, JavaScript, JSX, TSX, PHP, C#, Go, C, C++, Java, Kotlin, Rust and Perl. The README lists LZ4 compression, an in-memory SQLite database and fused Aho-Corasick pattern matching as the mechanisms behind the indexing speed, and states that memory is released after indexing completes.

The output is a persistent knowledge graph of functions, classes, call chains, HTTP routes and cross-service links. That graph is what the 15 MCP tools query. The README names search, trace, architecture, impact analysis, targeted index-coverage checks, Cypher queries, dead code detection, cross-service HTTP linking and ADR management among them. The binary also serves a 3D interactive visualization of the graph at localhost:9749 from the binary itself, which is a useful sanity check: you can look at what the indexer actually extracted before you trust a query result.

One design consequence is worth stating plainly. Hybrid LSP coverage is narrower than the language count. The README claims tree-sitter parsing across 161 languages, but semantic type resolution reaches a specific named list. For a language outside that list you get syntax-level structure, not resolved types, and queries that depend on type information will be weaker.

## Installing codebase-memory-mcp and indexing a first project

The README leads with a one-line install for macOS and Linux. It downloads and runs the install script from the main branch, so you are executing whatever is at that path at the moment you run it, not a pinned release artifact.

```bash
curl -fsSL https://raw.githubusercontent.com/DeusData/codebase-memory-mcp/main/install.sh | bash
```

On Windows the README gives a PowerShell sequence instead, and it explicitly recommends inspecting the script first. That advice is worth taking literally, because the installer modifies agent configuration files.

```powershell
Invoke-WebRequest -Uri https://raw.githubusercontent.com/DeusData/codebase-memory-mcp/main/install.ps1 -OutFile install.ps1
notepad install.ps1
Unblock-File .\install.ps1
.\install.ps1
```

If you prefer not to pipe a script into a shell, the README documents a manual path: download the archive for your platform from the latest release, extract it, and run the bundled installer. The archive names follow the pattern codebase-memory-mcp-<os>-<arch>.tar.gz for macOS and Linux, or .zip for Windows.

```bash
tar xzf codebase-memory-mcp-*.tar.gz
./install.sh
```

The installer accepts --skip-config for a binary-only install with no agent setup, and --dir=<path> for a custom location. After installing, restart your coding agent and tell it to index the project. The README's phrasing for the first real use is direct: say "Index this project". The install command auto-detects installed coding agents and configures them, and the README states it also strips macOS quarantine attributes and ad-hoc signs the binary, so no manual xattr or codesign step is needed.

## What the installer touches, and why that is the real limitation

The README is unusually candid here, and it deserves to be quoted rather than paraphrased: the tool "reads your codebase and writes to your agent configuration files. That is what it is designed to do." That single sentence is the adoption decision. If you are not comfortable with an installer that edits the configuration of every coding agent it detects on your machine, this is the wrong tool regardless of how fast the index is.

The README's mitigation is a release process rather than a sandbox. It states that for each release, three behaviourally identical executable candidates (unstripped, debug-stripped, stripped) are submitted to VirusTotal before testing, that the selected candidate is packaged with its SHA-256 unchanged, and that release notes link every measured candidate result. The project also documents a tolerance for a single Microsoft detection, `!ml`, in SECURITY.md. The README attributes that to a known false positive, noting that the same detection family hits gh, llama.cpp, Godot and Microsoft's own Go toolchain.

Read that as a process you can verify, not a guarantee. The verification it enables is specific: compare the SHA-256 of the archive you downloaded against the release notes, and read the linked VirusTotal result for the candidate you actually installed. The README also notes the full source is available if you want to audit before running. Separately, the README states all processing happens locally and code never leaves the machine, which is the claim to check if your threat model includes outbound network access.

The installer's breadth is a second, quieter limitation. It configures detected clients automatically and activates conditional clients only when their documented platform, marker or explicit existing config path is present. That is a careful rule, but it still means the set of files modified depends on what happens to be installed on your machine at install time.

## codebase-memory-mcp vs LSP-based and file-search approaches

The most useful comparison is against an LSP server, because both answer structural questions and they differ in what they persist. An LSP server holds a live, in-memory model of a project as you edit it, and its answers are current by construction. codebase-memory-mcp builds a graph ahead of time and serves queries from it. That makes it fast and cheap to query, and it makes staleness a real concern: an index built before your last refactor describes the code as it was. The README does not document an incremental re-index policy in the available documentation, so how you keep the graph current is something to establish on your own before relying on impact analysis in a long session.

The second comparison is against plain grep and file reads, which is what the README measures against. Grep has no index to go stale and no installer to trust. It also cannot answer a call-chain question without the agent reading a dozen files. The README's own framing puts the difference at roughly 412,000 tokens against roughly 3,400 for five structural queries, and cites a preprint, arXiv:2603.27277, reporting 83% answer quality, 10x fewer tokens and 2.1x fewer tool calls across 31 repositories. Those are the project's numbers from its own evaluation, not an independent result, and the preprint is where to check the methodology.

A third distinction is scope. Tools built around a single language and its compiler can resolve types deeply within that ecosystem. codebase-memory-mcp spreads across 161 languages for parsing and a named subset for semantic resolution, which buys breadth at the cost of depth outside that subset.

## Maintenance cost, licence and upgrade path

The last push to the default branch was on 2026-08-19, and the most recent release in the repository, v0.10.8, carries the same timestamp. Releases at 0.10.6, 0.10.7 and v0.10.8 landed on 2026-08-17, 2026-08-18 and 2026-08-19 respectively, so the version line is moving in small increments rather than long jumps. The repository is not archived.

Upgrade cost looks low on the surface. The binary is self-contained, grammars are vendored rather than fetched, and there is no language runtime, container or API key involved. That removes the usual class of breakage where a grammar or runtime version drifts out from under the tool. It also means a grammar update requires a new binary release, so a parsing bug in a language you care about waits on the project's release cadence rather than on you.

The licence is MIT, which permits commercial use, modification and redistribution with the licence text retained. That is a permissive position and it is not legal advice; if you vendor or redistribute the binary, read the LICENSE file and THIRD_PARTY.md, since vendored tree-sitter grammars carry their own terms. The repository also ships a DCO and a CONTRIBUTING.md, which is the governance surface to read if you intend to send patches.

## Conclusion

Adopt codebase-memory-mcp if you run a coding agent on a large repository and want structural queries answered from a prebuilt graph instead of file-by-file reads. Skip it if you cannot accept a tool that reads your source tree and writes to your agent configuration files, or if your work is mostly prose and configuration rather than code. Before trusting it, verify three things yourself: that the SHA-256 of the downloaded archive matches the release notes, that the VirusTotal candidate result linked in those notes is the one you installed, and that the agents it configured are the ones you actually use. The README states the full source is available for audit, and the MIT licence means you can read it before you run install.

## FAQ

### What is codebase-memory-mcp?

It is an MCP server that indexes a codebase into a persistent knowledge graph and exposes that graph to AI coding agents through 15 tools. It ships as a native executable for macOS, Linux and Windows, with no language runtime, hosted service or API key required.

### How do I install codebase-memory-mcp?

On macOS and Linux the README gives a one-line install script fetched from the main branch; on Windows it provides a PowerShell script that you download, inspect and unblock before running. There is also a manual path: download the platform archive from the latest release, extract it, and run the bundled installer. The installer supports --skip-config and --dir=<path>.

### Is codebase-memory-mcp safe?

The README states plainly that the tool reads your codebase and writes to your agent configuration files, and that all processing happens locally with code never leaving your machine. It also describes a release process where three executable candidates are submitted to VirusTotal, the selected candidate is packaged with its SHA-256 unchanged, and release notes link each measured result. The README recommends auditing the full source before running if you want to verify behaviour first.

### What are the differences between Graphify and Codebase memory MCP?

The README does not describe Graphify, so no comparison can be made from it. What can be said is how codebase-memory-mcp works: tree-sitter parsing across 161 languages with Hybrid LSP semantic type resolution for a named subset, producing a persistent graph queried through 15 MCP tools.

### What are the key differences between Codebase memory MCP and Codegraph?

The README does not cover Codegraph, so the difference cannot be stated from it. The relevant self-description is that codebase-memory-mcp builds a persistent knowledge graph of functions, classes, call chains, HTTP routes and cross-service links, and serves it to agents over MCP from a single static binary.

## Sources

- [Official documentation](https://deusdata.github.io/codebase-memory-mcp/)
- [Official README](https://github.com/DeusData/codebase-memory-mcp#readme)
- [Project repository](https://github.com/DeusData/codebase-memory-mcp)
- [Release notes](https://github.com/DeusData/codebase-memory-mcp/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/deusdata-codebase-memory-mcp
