# DiffMem stores agent memory in orphan branches, then shells out to Git to search it

> DiffMem is a Python memory service that keeps each user's notes as Markdown on an orphan Git branch and lets a retrieval agent grep the history. Its compose file defaults to no authentication and to empty model names, its second deployment chapter ends mid-sentence, and no release has ever been published.

**Growth-Kinetics/DiffMem** — Git Based Memory Storage for Conversational AI Agent

- Repository: https://github.com/Growth-Kinetics/DiffMem
- Stars: 905 · Forks: 61
- Language: Python
- License: not declared
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/growth-kinetics-diffmem

## Isolation is an orphan branch, not a separate repository

Every user lives in the same on-disk repository. The service allocates an orphan branch named `user/{user_id}` and checks it out into a per-user worktree when that user becomes active, so two branches carry no shared history at all and neither can see the other's commits. Storage splits into two pluggable halves. The storage backend decides where the repo and worktrees live, and the docs call it a hard requirement because the retrieval agent shells out to `grep` and `git log` against a real directory. The backup backend is optional and defaults to `none`, which means the only off-box copy is whatever a volume snapshot captures. Setting `BACKUP_BACKEND=github` with `GITHUB_REPO_URL` and `GITHUB_TOKEN` mirrors user branches into a private repository you own. Pushes run on a scheduler, `BACKUP_INTERVAL_MINUTES` defaulting to 30, and never sit in the request path. Pulls happen when a worktree is mounted, which the docs place at the first request per user after a restart.

## The retrieval agent has one tool and no index of its own

Retrieval does not run through a search library. The retrieval agent is a multi-turn LLM agent with a single `run(command="...")` tool that executes sandboxed shell commands against the memory repository. It reads `index.md` first, then reaches for `grep`, `git log`, `git diff` and `git blame`, and finishes by emitting a structured retrieval plan naming file sections, diffs and commit logs, which the service resolves into context. The wager is that current files stay small enough to reason over while history stays parked until something needs it, with no vector databases, no embeddings and no BM25. Its own roadmap concedes the weak seam: the indexing strategy carried over from the proof of concept is called too memory intensive without need, and context caps on retrieval are still waiting on a parametrized method. A second entry admits that an entity sometimes turns into a catch-all that keeps overloading itself with unrelated detail.

## The writer only ever touches the working tree

Writing follows the same shape. The writer agent analyzes conversation transcripts, identifies or creates entities, stages the resulting updates in Git's working tree, and commits them as explicit, atomic commits. That separation is the whole design: the checked out files hold the now view of a person, and the commit graph holds how that view moved. The reasoning behind it is token economy, since only the current surface enters the prompt and history is pulled selectively when a question calls for it. In production the project points at Annabelle, a simulated intelligence on WhatsApp and Messenger that it says maintains memory across thousands of conversations, referencing details from weeks earlier, tracking how a relationship shifts over time, building structured understanding of each person, and consolidating memories automatically as conversations grow. A companion repository shows the same pipeline working through a novel chapter by chapter.

## REQUIRE_AUTH defaults to false while the port is still published

Authentication is off unless you turn it on.

```
      REQUIRE_AUTH: ${REQUIRE_AUTH:-false}
      API_KEY: ${API_KEY:-}
```

The ports block publishes `"${PORT:-8000}:${PORT:-8000}"` on the host. The walkthrough argues you can leave the default when the only caller is another service on the same Coolify instance, and to set `REQUIRE_AUTH=true` plus a long random `API_KEY` when the domain is public. There is a tension worth naming before you follow either instruction. The same walkthrough credits Coolify with no open ports on the host, yet that ports mapping binds the port directly whenever the compose file runs as written. Treat the published port and the auth flag as one decision rather than two.

## Two model variables resolve to empty strings and fail late

The only variable the compose file marks required is `OPENROUTER_API_KEY`. The two model settings ship with empty defaults and a comment that spells out the consequence:

```
      DEFAULT_MODEL: ${DEFAULT_MODEL:-}  # required at runtime; set in Coolify or .env
      RETRIEVAL_MODEL: ${RETRIEVAL_MODEL:-}
      PORT: ${PORT:-8000}
```

Compose starts happily with either one blank, so a deployment can go green, pass the healthcheck, and still have no model to call. Nothing visible in the configuration validates those values at boot. The container also runs a single process, `uvicorn diffmem.server:app --host 0.0.0.0 --port ${PORT:-8000} --workers 1`, so there is no second worker to absorb the first failing request. Set both names in the Coolify environment tab, not just the API key, before pointing a client at the box.

## Two dependency manifests, and no image is ever published

Dependencies are declared twice with different strictness.

```
# DiffMem: Git-native memory backend
requests
openai
gitpython
python-dotenv
```

`pyproject.toml` pins ranges such as `requests = "^2.31.0"`, `gitpython = "^3.1.40"`, `openai = "^1.0.0"` and `python-dotenv = "^1.0.0"`, and marks `fastapi`, `uvicorn`, `pyjwt`, `httpx`, `python-multipart`, `aiofiles` and `hatchet-sdk` as optional extras. The list above carries no constraints at all, and the Dockerfile installs from a third file, `requirements-server.txt`. The version is `0.4.0` with a status of `Development Status :: 4 - Beta`, but the project has no GitHub releases, which is why `image: ghcr.io/growth-kinetics/diffmem:latest` sits commented out under a note that prebuilt images will appear once a release is tagged. Every deployment is a source build, gated by:

```
HEALTHCHECK --interval=30s --timeout=5s --start-period=10s --retries=3 \
    CMD curl -fsS http://localhost:${PORT:-8000}/health || exit 1
```

## Git is preconfigured globally, wildcard safe.directory included

The image configures Git before any repository exists, at global scope, with a fixed identity and an empty credential helper:

```
git config --global user.name "DiffMem" && \
git config --global user.email "diffmem@localhost" && \
git config --global credential.helper "" && \
git config --global http.postBuffer 524288000 && \
git config --global --add safe.directory '*'
```

That wildcard turns off Git's ownership check for every path in the container, which is convenient when worktrees are created by a different user id than the one running the shell, and it removes a guard that would otherwise fire on a mismatched owner. `tini` is installed and used as PID 1 so a SIGTERM from Coolify or Docker reaches the FastAPI lifespan shutdown, which the Dockerfile comment says flushes final backups. Reasonable choices for a single-purpose container, but worth knowing before you mount anything beyond the two state directories the image creates, `/data/storage` and `/data/worktrees`.

## The Hatchet chapter stops mid-sentence and MIT is only a classifier

Two parts of the project's own record are incomplete. The deployment guide reaches a heading, `### Production deployment with Hatchet`, and one line that begins `For durable, obser` and ends there. The rest of that chapter is not present, even though `hatchet-sdk = "^1.33.0"` sits in the optional dependency list and a `deploy/` directory sits at the root. Licensing is split the same way: the classifiers in `pyproject.toml` carry `License :: OSI Approved :: MIT License` and the badge at the top of the README points at the MIT text on opensource.org, yet no `LICENSE` file appears among the root entries, which run `.dockerignore`, `.env.example`, `.gitignore`, `Dockerfile`, `README.md`, `deploy/`, `docker-compose.yml`, `docs/`, `notes/`, `ontologies/`, `pyproject.toml`, `repo_guide.md`, `requirements-server.txt`, `requirements.txt`, `scripts/`, `src/`, `test_context.py` and `tests/`. Treat the MIT claim as a classifier and a badge until a file confirms it. The last push landed 2026-08-28.

## Conclusion

DiffMem is a sensible pick for a team that wants agent memory to stay as readable Markdown under Git instead of inside a proprietary store, and a poor fit for anyone who needs published binaries, a documented recovery path, or a hardened default posture. Before exposing it, set REQUIRE_AUTH=true with a real key, populate DEFAULT_MODEL and RETRIEVAL_MODEL so the first request is not what discovers they are empty, and turn on the GitHub backup backend if a local volume snapshot is not a backup plan you trust. The unfinished Hatchet section and the absent license file are worth resolving before this holds anything you cannot recreate.

## FAQ

### Does DiffMem need a vector database or embeddings?

No. Storage is Markdown files plus Git history, and the project states it uses no vector databases, no embeddings and no BM25. Retrieval runs through shell commands against a real directory, which is why the storage backend must be a mounted disk.

### How does DiffMem keep one user's memories away from another's?

Each user gets an orphan branch named `user/{user_id}` inside one local storage repository, checked out into a per-user worktree when active. Orphan branches share no history, so commits on one branch are unreachable from another.

### Which model provider does DiffMem call?

OpenRouter. `OPENROUTER_API_KEY` is the only variable marked required in docker-compose.yml, and `DEFAULT_MODEL` and `RETRIEVAL_MODEL` are read at runtime with empty defaults. The documentation does not enumerate which models are supported.

### Does DiffMem publish Docker images or GitHub releases?

Not yet. The repository has no GitHub releases, and the ghcr.io image line in docker-compose.yml is commented out with a note that prebuilt images will be published once a release is tagged, so the compose file builds from source every time.

### What is DiffMem's license?

The pyproject.toml classifiers list `License :: OSI Approved :: MIT License` and the README badge links to the MIT text on opensource.org, but no LICENSE file is listed at the repository root. Confirm the terms with the owner before relying on either.

## Sources

- [Growth-Kinetics/DiffMem on GitHub](https://github.com/Growth-Kinetics/DiffMem)
- [Issues](https://github.com/Growth-Kinetics/DiffMem/issues)
- [README](https://github.com/Growth-Kinetics/DiffMem/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/growth-kinetics-diffmem
