# hippo-memory's decay tied with decay switched off

> A SQLite memory store for coding agents with zero runtime dependencies, BM25 retrieval by default and hooks for the major agent CLIs. Its own receipts page reports that decay made no difference, that sleep lowered recall, and that one benchmark claim was retracted.

**kitfunso/hippo-memory** — Biologically-inspired memory for AI agents. Decay, retrieval strengthening, consolidation. Zero dependencies.

- Repository: https://github.com/kitfunso/hippo-memory
- Website: https://hippo-memory.com
- Stars: 770 · Forks: 44
- Language: TypeScript
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/kitfunso-hippo-memory

## Decay tied, sleep hurt, and only two mechanisms measured well

The pitch is forgetting: good memory is knowing what to forget, so a memory marked wrong drops out of the top results and unused ones fade. Then the same README says decay tied with decay switched off, and that sleep lowered recall.

So the borrowing from the hippocampus, decay plus three layers plus sleep consolidation, is described as inspiration and not as the working part. Two other mechanisms are the ones the author says were measured helping: mark a recalled memory bad and it ranks lower, and memories you keep using get stronger. Those two were checked on a synthetic test, which is a smaller claim than the headline suggests.

The third mechanism, superseding a changed fact with `hippo supersede` so the old version leaves recall, carries an explicit note that it has not been measured. Storage, then, is the part with evidence behind it, and forgetting is the part still being argued.

## hippo sleep sends memory text to Anthropic by default

Recall makes no network call. Sleep does, and that is the default once a key exists. `hippo sleep` sends memory text to Anthropic for fact extraction when `ANTHROPIC_API_KEY` is set. Sleep runs at the end of every Claude Code and OpenCode session and inside the daily job, so with the key present the call happens without being asked for. The switch is a config key:

```json
{"extraction":{"enabled":false}}
```

The opt-in extras behave the same way and are named: the Jev reranker, the LLM reranker and the API embedders all call out.

The Slack connector side is the opposite case and the project proves it rather than asserting it. A 1000-event ingestion smoke run reports zero outbound HTTP, and the evidence is a `globalThis.fetch` spy that throws if called, not a hardcoded zero. That is the standard worth holding other claims to.

## Two LongMemEval numbers that are not a before and after

The headline figure is BM25 only, no embeddings, all 500 questions pooled into one store on the LongMemEval oracle split at v0.11: R at 5 of 74.0 percent. The README then points at a second set of LongMemEval results under its benchmarks section and says explicitly that this is a different setup, so the two are not a before and after.

That is the kind of sentence most projects leave out. Pooling every question into one store is the easy case for any retriever, and per-haystack numbers are the harder one. Reading 74.0 as an improvement claim would mean reading the wrong comparison.

The other receipts are narrower and more honest about what they are. A staged Slack corpus of ten incident scenarios had recall beat transcript replay in ten of ten, with the answers sitting mid-channel by design: the top 10 results held every answer message while the channel's last 10 held none. And the test suite is described as 3,500 or more tests on a real database, with no module mocks and no mocked store, only paid network calls stubbed.

## The paid reranker wins on ranking and loses on answers

The TypeSafe Jev reranker is off by default and costs about 0.0004 USD a recall. On a private 300-query developer store it moved R at 1 from 0.41 to 0.62 with `hippo recall "<query>" --reranker jev`. The statistics are given rather than implied: a 2000-draw paired bootstrap, a margin that held in 20 of 20 seeds, and a permutation null that reached it in 0 of 200 runs.

Then the negative half. Three graded tests on one 150-question LongMemEval set did not show a better answer rate than the free local cross-encoder, and that result sits in the same document. What the money buys today is a shorter context: on that set, two memories ranked by Jev answered as well as five ranked by the cross-encoder. Ranking is not answering, and the gap between those two is the whole argument for paying per recall.

The local alternative costs nothing to run: bring your own Transformers.js with `@huggingface/transformers`, or the legacy `@xenova/transformers`. Nothing is auto-installed, for embeddings or anything else.

## Installing the package enables nothing on its own

The whole install is one command:

```bash
npm install -g hippo-memory && hippo init
```

The README is explicit that package installation alone does not enable automatic preservation on every agent, and that coverage depends on the integration and on host trust. What `hippo init` does differs per agent, and the differences matter:

- Claude Code and OpenCode get hooks, Claude Code's in `~/.claude/settings.json`, plus a daily 6:15am run.
- Codex gets two hooks in its `hooks.json`, and Codex runs them only once you trust them in `/hooks`.
- Codex, Cursor, OpenClaw and Pi get instructions added to an `AGENTS.md` that already exists, which means a project without one gets no file created.

Init is per project. Covering every git repo under a folder is a second documented step, and any MCP client can connect over HTTP instead.

## Zero runtime deps, Node 22.16 or later, and no named SQLite binding

The storage line is short and specific: a SQLite backbone with markdown mirrors, git-trackable and human-readable, so the store can be committed and read without the tool. Search is BM25 out of the box, with no model and no network call, and embeddings are optional.

The dependency claim is zero runtime deps with a floor of Node.js 22.16. Those two facts together mean the SQLite layer cannot be coming from npm, and the README never names which implementation it uses. It also lists what it imports from, which is the portability argument: ChatGPT exports, CLAUDE.md, Cursor `.cursorrules`, Slack history and plain markdown all end up in one store.

Inside the repository the picture is busier than the claim. There is a `ui/` directory with its own build script, its own npm install and its own `npm audit` pass, a `python/` directory, a `website/`, a `deploy/` directory with a `.dockerignore` and no Dockerfile, and research files at the top level: RESEARCH.md, ROADMAP-RESEARCH.md, TODOS.md, PLAN.md, CONTEXT.md and MEMORY_ENVELOPE.md.

## The published package drops dist/src and every sourcemap

The files list in package.json is where the release shape shows:

```json
"files": [
  "dist",
  "!dist/src",
  "!dist/**/*.map",
  "dist-ui",
  "bin",
  "scripts/postinstall.cjs",
  "openclaw.plugin.json",
  "extensions/openclaw-plugin"
]
```

The main entry is compiled, `./dist/index.js`, and the tarball then excludes the source subtree and every source map from it. Stack traces from the published package will not point back to TypeScript lines.

The OpenClaw extension is the opposite: it ships as `index.ts` source, so that part of the package is uncompiled by design. The build runs three separate TypeScript passes, the main one plus `tsconfig.benchmarks.json` and `tsconfig.extensions.json`, and there are two smoke scripts, `smoke:pack` and `smoke:openclaw-install`, that check the packed result rather than the working tree.

## A script that checks release notes for em dashes

The release machinery in package.json is unusually specific. The `version` script synchronizes versions and then checks manifests. `prepublishOnly` runs a manifest version check, a changelog fragment check, `check-em-dashes-in-release-notes.mjs` and another script the excerpt cuts off mid-name. Fragments live in `changelog.d/` rather than being written straight into CHANGELOG.md.

Em dashes in release notes are machine-rejected. That is a writing rule encoded as a gate, and it is enforced at publish time rather than by review.

The release cadence is the other signal. v1.52.8 shipped on 2026-09-28 to stop terminal popups when a session closes on Windows, v1.52.9 followed on 2026-09-29, and v1.53.0 landed on 2026-09-30, matching the version in package.json. Three releases in three days on a default branch called master, with a retracted benchmark claim from v1.7.9 still referenced in the README, means the tag you install and the behaviour described in the text can be a release apart.

## Conclusion

hippo-memory is worth reading for its receipts page alone: a retracted claim, a negative result kept next to a positive one, and two LongMemEval numbers the author refuses to compare. The mechanism that carries the weight is not decay but marking a memory bad and letting repeated use strengthen what survives, and even the superseded-fact path is marked unmeasured. Two operational points come before any feature evaluation. First, installation is not configuration: `npm install -g hippo-memory` does nothing on its own, and preservation depends on `hippo init` plus host trust, which differs per agent. Second, with `ANTHROPIC_API_KEY` set, `hippo sleep` sends memory text to Anthropic at the end of every Claude Code and OpenCode session and in the daily 6:15am job, and `{"extraction":{"enabled":false}}` in `.hippo/config.json` is what turns that off. Three releases landed in the three days to 2026-09-30, so pin a version before wiring it into anything you cannot easily undo.

## FAQ

### Does hippo-memory need an API key?

Not for search. Retrieval is BM25 with no model and no network call, and embeddings are an optional install you bring yourself. The exception is `hippo sleep`, which sends memory text to Anthropic for fact extraction when ANTHROPIC_API_KEY is set, at the end of every Claude Code and OpenCode session and in the daily job.

### Does installing the hippo-memory package turn on automatic memory saving?

No. Package installation alone does not enable automatic preservation on every agent. Run `hippo init` inside a project, complete the documented setup and the host trust each agent requires, and accept that capture and compaction coverage depend on the integration.

### Which coding agents does hippo-memory hook into?

Claude Code and OpenCode receive hooks, Codex receives two hooks in its hooks.json that you must trust in /hooks, and Codex, Cursor, OpenClaw and Pi receive instructions added to an existing AGENTS.md. Any MCP client can connect instead.

### Did hippo-memory retract a benchmark claim?

Yes. The informal magnitude claim from the Sequential Learning Benchmark at v0.11.0 was retracted at v1.7.9, with the mechanism still shipped. The CHANGELOG entry for v1.7.9 is the reference, and failed runs are kept next to the results in docs/evals.

### Does decay improve recall in hippo-memory's own results?

No. The README states that decay tied with decay switched off and that sleep lowered recall. What it reports as measured helping is marking a recalled memory bad so it ranks lower, and repeated use strengthening a memory, checked on a synthetic test.

## Sources

- [kitfunso/hippo-memory on GitHub](https://github.com/kitfunso/hippo-memory)
- [License: MIT](https://github.com/kitfunso/hippo-memory/blob/master/LICENSE)
- [Project website](https://hippo-memory.com)
- [README](https://github.com/kitfunso/hippo-memory/blob/master/README.md)
- [Releases](https://github.com/kitfunso/hippo-memory/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/kitfunso-hippo-memory
