# Ix indexes your repository with tree-sitter and ships its database as an image nobody can rebuild

> A persistent graph of symbols, calls and imports, queried by a CLI, an MCP server and a visualizer, all reading one local ArangoDB backend. The open half is the indexer and the clients. The half that stores your graph is released as a Docker image that is explicitly not built from this repository.

**ix-infrastructure/Ix** — Understand any codebase instantly. System intelligence for codebases, built for humans and AI.

- Repository: https://github.com/ix-infrastructure/Ix
- Website: https://www.ix-infra.com
- Stars: 1,022 · Forks: 79
- Language: TypeScript
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/ix-infrastructure-ix

## The root manifest reads 0.5.0 and describes a different product

The repository's root manifest is four fields long: a name, a version, a private flag and a description. There are no scripts, no dependencies and no entry point.

The version it carries is 0.5.0. The newest published release is 0.12.0, with 0.11.1 and 0.11.0 behind it in the same week. So the root manifest trails the release line by seven minor versions and moves on a different schedule from the tags, which is what you would expect if it were a stub rather than the published artifact's manifest.

The description is more confusing. It calls the project Ix Memory and describes persistent, time-aware context for language model assistants. The project described in the file itself is Ix, which parses a repository into a graph of symbols, calls and imports for coding agents. A memory product and a codebase indexer are different products, and the root manifest names the first one.

None of this breaks anything, because the root manifest is not what installs the CLI. The tree has separate directories for the command line client, the ingestion layer, the documentation site, the Homebrew formula and the tap. Someone auditing versions should read the manifest in the command line directory rather than the one at the root, and anyone looking for what this software is should read the file rather than the package description.

## The backend is a released image, and the file says it is not built here

The architecture is a parser, a graph, and three clients, and one of those three layers is closed to a source build.

The parser is tree-sitter, extracting symbols, calls and imports. Those become nodes and edges describing what calls what, contains what and imports what. The graph is stored in a local backend made of ArangoDB plus a memory layer, run for you in Docker.

The sentence that matters is short and explicit: the backend ships as a released Docker image, and it is not built from this repository.

So the part of the system that holds your code's structure is an image you pull rather than a program you build. Everything a reader would normally check is available for the part that produces the data, and unavailable for the part that keeps it. The clients, the HTTP endpoints all three share, and the ingestion code are open under the Apache licence, and the store is not.

That is a legitimate way to ship a database-backed service, and it is also the kind of thing that should change what you are willing to point at proprietary code. The repository does have a standalone compose file at its root for running the released image on its own, so the deployment is reproducible even though the build is not.

## The headline token saving runs from 30 to 99.7 percent, and disclaims itself

The results section is four sentences and two of them are a disclaimer.

The claim is that, across the company's own development work, querying the graph instead of feeding files into the prompt cut token use by between thirty and ninety-nine point seven percent, varying widely with the task and the size of the repository. The next sentence says these are internal measurements and not a published benchmark.

Both halves are true and they belong together. A range spanning almost the whole interval from 30% to 100% is not a measurement of the technique, it is a measurement of the spread between the easiest and hardest task in one organisation's own work. There is no dataset named, no task list, no model named, no prompt shape, and no definition of which tokens were counted.

What follows the claim is the better argument, and it does not need a number. Explaining a symbol returns that symbol and its immediate relationships. Answering the same question by reading the file and the files it imports costs more once, and costs it again in the next session, because nothing about the previous reading survives.

That argument is about bounded retrieval rather than compression, and it holds regardless of what the percentage turns out to be. A reader should take the mechanism and leave the figure.

## Intel Mac users compile the CLI, and the installer brings its own tools

The installer is a single pipe to a shell on macOS and Linux, with a PowerShell equivalent for Windows, and it is more than a download.

It checks for and installs what is missing: a supported Node.js version, Git, ripgrep, which is what powers the text search subcommand, and Docker plus Docker Compose for the local backend. The stated prerequisite is only a terminal with a download tool. On Windows the requirements are narrower in one way and wider in another, since Node and Docker Desktop have to be in place first.

The pre-built packages cover Apple Silicon macOS, Linux on both x86-64 and arm64, and Windows on x86-64. Intel Macs have no pre-built package at all, and the documented route is a Homebrew tap that builds from source.

That gap is the practical cost of the packaging matrix. An Intel Mac user is compiling a TypeScript project on first install, which means a working local toolchain, and the file does not say how long that takes or what can go wrong.

The release cadence on the other side of that is fast. Three of the last releases are 0.11.0, 0.11.1 and 0.12.0 across four days at the end of September and the first of October, so the tool is moving quickly enough that a from-source build is a reasonable alternative to a lagging binary.

## The MCP installer refuses to overwrite a name, and its own name is ambiguous

Registering an editor is one command, and the safety properties are stated rather than assumed.

The installer detects which clients are present and registers the server with each. It supports seven of them, and it writes through each client's own MCP command where one exists. The guarantee is that it never overwrites a server name it does not own, with a force flag to replace one deliberately and a host flag to limit which client is touched. There is a dry run that reports what would change and writes nothing, and a doctor command to check what is currently registered.

The guarantee rests on knowing which name is yours, and the file uses two. The installer is described as registering the MCP server under the name of that command, while the manual example for one client adds the server under a different name, the memory one, and the Claude Code native plugin is installed under that same memory name.

So a reader cannot tell from the file whether the tool considers itself to own the command name or the memory name, and which of your existing registrations it will leave alone depends on that answer. It is a small thing to get wrong and an easy one to check by running the dry run first, which is exactly why the dry run exists.

The per-client write mechanics, repair steps and process isolation are deferred to a separate document, which is the right place for them.

## Seven install paths for one tool, and six of them are other repositories

The number of ways to get this in front of an editor is the most surprising number in the file.

There is the main installer with a PowerShell variant, the MCP installer that handles seven clients on its own, a manual command for registering a single client by hand, and a shell script that deploys an agent skill to every harness it can find, again with a dry run.

Then there are the native plugins, and these live in six sibling repositories rather than in this one: one for Claude Code, added through its own marketplace command and installed under the memory name; one script for Codex, with a PowerShell version at the same address; a plugin package for OpenClaw installed through that tool's own plugin command; a Gemini extension installed from its repository URL; and shell installers for OpenCode and for Cursor that pipe a remote script into the shell.

So the surface area a user can touch is this repository, six others, and a set of pipe-to-shell scripts that need auditing before running, which the file does address for the main installer by pointing at it and for the others not at all.

The agent skill is the tidiest of the options. One skill file follows both the Claude Code skill format and a cross-harness agents standard, so four named clients and any agent reading the shared skills directory load the same tree from their own location. That is a single file to review, which is more than can be said for six installers.

## Three clients read one graph, and the first section sells a different product

The durable part of the design is small: one graph, three ways to read it.

The command line client answers directly. The MCP server is what an agent connects to. The visualizer opens from a view command. All three read the same backend over the same HTTP endpoints, which are documented in one place, so a feature added for one client is available to the others by construction.

The four commands the file leads with are the interface worth memorising:

```bash
ix map .
```

```bash
ix explain AuthService
```

```bash
ix trace user_login_flow
```

```bash
ix impact verify_token
```

That is build the graph for a directory, explain what a symbol is and what it touches, trace how a named flow actually runs, and show what breaks if a given symbol changes. The stated property of the first three is that they return bounded answers about one symbol or flow rather than whole files.

The graph persists locally between sessions and agent runs, which is the claim the whole comparison table rests on: architecture mapped once and kept, rather than re-derived every session.

The first section of the file, above all of this, is about a different product. Kartr is described as an agent platform on the same memory engine, extended to the sources a codebase does not contain: documents and files, email and calendar, meetings and notes, and team chat. It is in alpha and onboarding early users through a web form. Worth knowing that the top of the file is a cross-sell for an alpha service, and that the memory engine the two share is the one shipped as an image.

## Conclusion

Ix fits a team whose agents keep re-reading the same files to answer the same architecture questions, and where the graph can live on a machine you control. The query model is the right idea: a symbol's immediate relationships instead of its file, a call flow instead of a directory listing, and a graph that survives a session. Four things to check before committing to it. That the storage layer is a released image rather than a build from the repository, so the part that holds your code's structure is the part you cannot audit from the source. That you are on a platform with a pre-built CLI, since Intel Macs compile from source. That you can live with Docker plus a local database process on the machine, since the installer wants both. And that you have a way to judge the token saving yourself, because the headline figure is explicitly internal with no methodology attached, and a range that wide could be one easy task or a thousand.

## FAQ

### What does Ix do to a repository?

It parses the repository with tree-sitter, extracts symbols, calls and imports, and persists them as a local graph of what calls, contains and imports what. Three clients read that graph: the command line, an MCP server, and a visualizer, all sharing one HTTP API, so an agent can ask about a single symbol or flow instead of receiving whole files.

### What are the four main Ix commands?

The map command builds the graph for a repository, the explain command returns a symbol and its immediate relationships, the trace command follows a named flow, and the impact command shows what breaks if a symbol changes. Results are stored locally and persist between sessions and agent runs.

### How much token use does Ix actually save?

The README quotes a reduction of 30 to 99.7 percent from internal measurements across its own development work, and states in the same paragraph that these are internal measurements rather than a published benchmark, varying widely with the task and the size of the repository. No dataset, task list or counting method is given.

### Which AI clients does the Ix MCP installer support?

Claude Code, Codex, Cursor, VS Code, Gemini CLI, OpenClaw and opencode. It writes through each client's own MCP command where one exists, never overwrites a server name it does not own, and offers a dry run that writes nothing, a force flag, a host flag to limit the run, and a doctor command to check registrations.

### Does Ix have a pre-built package for Intel Macs?

No. Pre-built CLI packages cover Apple Silicon macOS, Linux on x86-64 and arm64, and Windows x86-64. For an Intel Mac the documented route is a Homebrew tap that builds the CLI from source, and the installer otherwise brings in Node.js 22+, Git, ripgrep and Docker with Compose.

## Sources

- [ix-infrastructure/Ix on GitHub](https://github.com/ix-infrastructure/Ix)
- [License: Apache-2.0](https://github.com/ix-infrastructure/Ix/blob/main/LICENSE)
- [Project website](https://www.ix-infra.com)
- [README](https://github.com/ix-infrastructure/Ix/blob/main/README.md)
- [Releases](https://github.com/ix-infrastructure/Ix/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ix-infrastructure-ix
