# iai-pme: a local memory engine for MCP coding agents

> iai-pme is a Python and Rust MCP server that captures every turn of a coding session verbatim and injects a small memory pack at session start. It is local, MIT licensed, and aimed at one person's assistant rather than a product you ship.

**CodeAbra/iai-personal-memory-engine** — A cyber brain for your AI. It never forgets a detail, remembers exactly what you said, and learns how you work over time. Free, local, works with Cursor, Claude Code, Codex, OpenClaw, Hermes and more. MIT.

- Repository: https://github.com/CodeAbra/iai-personal-memory-engine
- Stars: 894 · Forks: 110
- Language: Python
- License: MIT
- Published: 2026-08-08 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/codeabra-iai-personal-memory-engine

## What iai-pme solves, and for whom

Every new chat with a coding assistant starts empty. You re-explain the repo layout, the naming convention you settled on last week, the reason one module was split. The README frames the target user plainly: a single person who wants their own assistant to remember them, not a team building a memory feature into a product.

The distinction matters because the README draws it itself. Mem0, Supermemory, Graphiti and Letta are described as memory layers for products you build, cloud-first, with LLM extraction pipelines. claude-mem is described as compression-based, storing AI-generated summaries of sessions. iai-pme takes a third position: verbatim capture with no LLM in the write path. According to the README, a memory write costs file IO plus a local embedding, while an extraction-based layer costs an LLM call per fact.

That choice has consequences the README leans into. Because turns are stored write-once and never rewritten, a fact that later changes leaves the old version archived and still retrievable. The project reports Rescue@10 of 1.000 and historical wording of 1.000 on its own benchmarks. Those are the author's numbers from the harness shipped in bench/, not an independent result.

The intended user is also the person who wrote it. The README says the author built it for himself, has been running it daily for months, and released it after the fact. That origin explains both the strengths and the gaps: the design reflects one workflow, and the documentation of edge cases reflects one person's experience of them.

## Verbatim capture, then a map assembled over time

The mechanism has three stages. Capture hooks record every turn of every session as it happens. A background daemon then organizes those captures into what the README calls a personal map. At the start of a new conversation, the MCP server serves a small slice of relevant memory back to the host.

The organizing step is where the project departs from a wrapper around an existing stack. The README states that the storage engine, the community-detection algorithm, the hyperdimensional memory substrate and a native engine are its own code. The Cargo workspace lists eight members, including crates/lilli-hd, crates/lillibrain, crates/lilliengine and crates/lilli-parity, plus four rust/ crates for embedding, graph and vector cores.

The Python side has been thinned out over time. In pyproject.toml, networkx is described as evicted from the main dependencies; the graph algorithm surface routes through a native Rust namespace, and the legacy MemoryGraph backend lazy-imports networkx inside __init__ as an internal detail pinned to networkx==3.3 for tests. numba is listed as JIT for the community-detection kernels, scipy for CSR build glue. The practical reading: this is a two-language project, and the parts doing the heavy work are compiled.

The Python dependencies are still substantial, and they hint at what the daemon does per session. cryptography provides AES-256-GCM at rest, keyring stores keys in the OS keychain on macOS, Linux Secret Service or Windows Credential Manager, cachetools supplies a TTLCache used as an activation-cascade LRU warmer, and psutil feeds a self-process CPU watchdog in the daemon. pandas is named as the Hippo query-result format behind a .to_pandas() reader. None of that is optional decoration; each entry maps to a running component.

The README also notes that the anthropic SDK was removed at v7.5, replaced by a subscription-billed claude -p subprocess rather than direct calls to api.anthropic.com. That is a meaningful design decision for anyone evaluating cost, because it moves the model call onto whatever subscription the CLI already uses.

## Installing iai-pme and getting a first memory pack

The README gives a single bootstrap command for macOS and Linux. It checks prerequisites, clones the repo to ~/.local/share/iai-pme, builds, installs the background engine and the capture hooks, and registers the MCP server.

```bash
curl -fsSL https://raw.githubusercontent.com/CodeAbra/iai-personal-memory-engine/main/scripts/bootstrap.sh | bash
```

Piping a remote script into bash is the project's own recommended path, and it is worth reading scripts/bootstrap.sh before running it. The README does not document a rollback command for the bootstrap.

The package itself is on PyPI as iai-pme, and pyproject.toml requires Python >=3.11,<3.13, so a 3.13 interpreter will not satisfy the constraint. Dependencies include numpy, scipy, numba, tiktoken, cryptography, keyring, cachetools, psutil and pandas.

```bash
pip install iai-pme
```

That installs the Python package. Note the caveat in setup.py: editable installs skip the npm build of the TypeScript MCP wrapper entirely, and the resolver falls back to mcp-wrapper/dist/ on an editable install. A plain pip install of the wheel compiles the wrapper during the build instead.

Once the server is registered with your MCP host, the first real check is the doctor command, which the README lists in its table of contents. The dashboard is the second check: it counts memory packs served, tokens injected, and a lower-bound estimate of tokens saved, measured against the search side rather than assumed.

If you would rather build from a checkout than run the bootstrap, the repository layout is the map. scripts/ holds the install scripts, mcp-wrapper/ holds the TypeScript MCP server, rust/ and crates/ hold the native workspace, and src/ holds the Python package. The Cargo.toml at the root is the workspace manifest, and rust-toolchain.toml pins the toolchain the build expects.

## The 88% token claim and what it actually covers

The README's headline number is that an injected memory pack costs roughly 88% less than the agent search it displaces. The arithmetic is spelled out: a ~350-token pack against a ~2,850-token search round-trip, the latter being a measured 2,639 tokens plus per-call overhead. The measurement comes from the author's own store over three recent weeks: 282 memory packs served, about 99,000 tokens injected, a lower-bound saving of about 707,000 tokens.

Two things narrow the claim. First, it is a lower bound by the engine's own conservative formula, and it comes from one person's usage, not a controlled comparison across projects. Second, the README explicitly excludes explicit mid-session recall from the figure. A memory_recall call is bounded by budget_tokens, default 1,500, and typically returns more than an ambient pack, so it saves against a search but not 88% of it.

That kind of self-limiting note is unusual in a README and makes the number easier to reason about. The dashboard computes your own count from your own store, which is the only figure that should drive a decision here.

There is a second framing worth separating from the first. The README calls the engine a local context provider for your MCP host, which is a different claim from saving tokens. A context provider replaces re-reading files the assistant already read; the token figure is the accounting of that substitution. The two are consistent, but the token number is downstream of the design, not the reason for it.

## Where iai-pme is the wrong tool

If you are shipping an application with many users, this is not the right layer. The README says so directly: use Mem0 or one of the others if you need multi-tenant memory for a product. iai-pme assumes one machine, one person, one local store with no account and no API key.

The operational cost is real too. A background daemon runs continuously, capture hooks are installed into your assistant's session lifecycle, and the native extension is built from a Rust workspace. The build is platform-conditional: setup.py adds an accelerate feature only on darwin, because the embedder's matmuls must link Apple Accelerate BLAS on macOS, and accelerate-src does not build on Linux. On macOS without that link, the comment says a 512-token encode runs a naive matmul and takes seconds instead of milliseconds. That is a build-environment dependency you inherit, not a runtime setting you can flip later.

The Python version pin is another boundary. requires-python is >=3.11,<3.13, so an environment already on 3.13 needs a separate interpreter before the package will install at all.

Finally, the project has no homepage, and the README does not document a rollback path for the bootstrap script. If your policy requires a documented uninstall procedure before you install anything, that procedure is not in the README.

## How it differs from a compression-based memory tool

The cleanest comparison in the README is against claude-mem, which it characterizes as storing AI-generated summaries of your sessions. The difference is not storage format, it is who does the summarizing and when.

With a summary-based approach, an LLM decides what was important at capture time, and the original wording is gone. With iai-pme, the turn is written verbatim and never rewritten; the organizing happens afterward, into a map, and the old version of a changed fact remains archived and retrievable. The README's stated memory style is verbatim over paraphrase, precise cues, rare events kept rare, which it calls autistic by design and explains under the About the name section.

The trade-off is storage growth and retrieval work. Keeping everything verbatim means the store keeps growing and the engine has to select a small slice at session start rather than reading a compact profile. The README's answer is the community-detection and hyperdimensional machinery, plus the native engine for speed. Whether that holds at a store far larger than the author's is not something the README addresses.

Against the cloud layers the difference is sharper still. Mem0, Supermemory, Graphiti and Letta are built for products, so they assume managed infrastructure, external vector databases or graph stores such as Qdrant, Neo4j or Postgres, and per-extraction LLM calls. iai-pme ships its own storage engine in Rust so there is nothing external to install, and the README lists that as a row in its comparison table.

## Licence, maintenance and upgrade cost

The project is MIT licensed, and the README states that everything is in the repo. The Rust workspace declares its crates as MIT OR Apache-2.0, which is a different identifier from the repository licence, so if you redistribute a built artifact you should read LICENSE and NOTICE.md rather than assume the two match. numba is noted in pyproject.toml as BSD-2, described there as MIT-compatible. This is a description of what the files say, not legal advice.

The last push was on 2026-08-23, and v3.0.8 carries the same timestamp. Releases v3.0.7 and v3.0.6 landed two days earlier, so the project was moving quickly in late August. The repository is not archived.

Upgrade cost comes from the two-language split. A version bump can pull a new Rust workspace dependency set (candle-core 0.10.2, tokenizers 0.23.1, pyo3 0.29, and others pinned in Cargo.toml) and rebuild the native extension, which is the slow part of any install. The README's Staying up to date section is the place to check what an upgrade actually does to an existing store.

One detail in Cargo.toml is worth knowing before you plan a build pipeline. hf-hub is declared with default-features off and only the ureq feature, specifically to keep native-tls, openssl-sys, reqwest and tokio out of the dependency tree so manylinux wheel builds need no OpenSSL headers. That is a deliberate constraint on the build, and changing it would change what a Linux build requires.

## What to check before you commit

Start with the doctor command the README lists, then open the dashboard and watch the counters for a week of real work. The token accounting there is computed from your store and your searches, which is more useful than the author's three-week sample.

Then test the thing the project is actually claiming: change your mind about something mid-project, and confirm that the old wording is still retrievable alongside the new. That is the behaviour the verbatim design is supposed to buy, and it is the one a summary-based tool cannot reproduce. If the archived version is not there, the design premise has not held on your data, and the token savings stop mattering.

Also check the store's growth over that week. Verbatim capture has no pruning step in the capture path by design, so the size of the store after a few hundred sessions is the number that tells you whether the local approach stays practical on your disk. The README does not publish a figure for that, which is exactly why it is worth measuring on your own machine.

## Conclusion

Adopt iai-pme if you already run an MCP host such as Claude Code or Cursor and want one person's session history to persist on your own disk. Do not adopt it if you need multi-tenant memory for an application you ship, or if you are unwilling to run a background daemon and a Rust build. Before trusting it, run the doctor command and check the dashboard's own token accounting against your own store, since the 88% figure comes from the author's measurements.

## FAQ

### What does iai-pme need to run?

It runs as a local MCP server over stdio, with a Python package requiring Python >=3.11,<3.13 and a Rust-built native extension. The README's quick start installs a background engine and capture hooks alongside the MCP server on macOS or Linux.

### Does iai-pme send my data anywhere?

The README states there is no API key, no account and no telemetry, and that the engine, the store and the embeddings all run locally. It says the only thing leaving your machine is the normal model call your CLI already makes.

### Which assistants does iai-pme work with?

The README lists Claude Code, Claude Desktop, Cursor, Codex CLI, Gemini CLI, Cline, Continue.dev, Zed, Cherry Studio, Goose, Aider, Hermes, OpenClaw, Le Chat and Kimi. Anything that speaks MCP over stdio is the stated requirement.

### Can I use iai-pme for an app I am shipping to many users?

The README advises against it and points to Mem0, Supermemory, Graphiti or Letta instead. iai-pme is described as a personal memory engine for the assistant you already use, assuming one machine and one local store.

## Sources

- [Official README](https://github.com/CodeAbra/iai-personal-memory-engine#readme)
- [Project repository](https://github.com/CodeAbra/iai-personal-memory-engine)
- [Release notes](https://github.com/CodeAbra/iai-personal-memory-engine/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/codeabra-iai-personal-memory-engine
