Model or dataset
Goldentrii/AgentRecall-X avatar
Goldentrii/AgentRecall-X

AgentRecall-X: a corrections ledger that measures whether your agent stops repeating a mistake

Correction-first persistent memory for AI agents. MCP server + SDK + CLI. Compounds across sessions.

371 stars60 forksJavaScriptMIT

At a glance

What is it?
AgentRecall-X is a correction-first memory layer for Claude Code and other MCP clients, published as an MCP server, SDK and CLI. Its own published numbers show why the measurement harness, not the memory store, is the interesting part.
Who is it for?
Adopt AgentRecall-X if you run Claude Code or another MCP client across many sessions and you want corrections stored as structured records with a severity and an outcome, not as free-form notes. Do not adopt it if you need a retrieval benchmark winner, a hosted service, or a memory layer that works without a client that speaks MCP.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 3 days ago.
What is it written in?
Mainly JavaScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem AgentRecall-X targets: corrections that vanish when the session ends

Most agent memory tools store facts the agent decided were worth keeping. AgentRecall-X stores the opposite: the moments where the human said the agent was wrong. The README frames this as a governed corrections ledger, where each correction is a structured record carrying severity, evidence and outcome tracking rather than a line of prose appended to a notes file. That record is meant to persist across sessions, across projects and across agent restarts.

The audience is narrow and specific. The README's install instructions are written for Claude Code first, with a generic MCP JSON block for other clients. There is no hosted dashboard and no cloud dependency in the default configuration; the badge row states cloud zero by default. If you are running a single short-lived chat and never correct the model twice about the same thing, the ledger has nothing to compound and the rest of the system is overhead.

Inside the loop: session_start, remember, session_end

The mechanism is a session lifecycle wrapped around a store. At the start of a session the agent calls session_start to load context. When the human corrects it, the agent calls remember with type "correction". At the end of a session it calls session_end to compound what it learned. The README gives exactly those three instructions as the first message of every new session.

Each correction accumulates a retrieved_count, and when the agent meets the same situation again the outcome is recorded as heeded or recurred. That second recording is the part the project claims nobody else does. The README states that LongMemEval, LoCoMo, MemoryAgentBench and the Letta Leaderboard all test retrieval or within-session updates, and that no public benchmark measures whether a captured correction changes what a fresh agent does in a new session.

The repository is a workspace, not a single package. package.json declares packages/core, packages/mcp-server, packages/sdk and packages/cli, and the build script compiles them in that order. The Dockerfile builds only the MCP server workspace and starts packages/mcp-server/dist/index.js. Retrieval is described in the badge row as keyword plus RRF, so there is no embedding model in the default path. The README reports a median session_start injection of 1,489 tokens and a p95 warm latency of 363 ms, both attributed to UPDATE-LOG.md section C2.

Installing the MCP server and running a first correction

The README gives the Claude Code install as a single command, scoped to the user so the server is available across projects:

bash
claude mcp add --scope user agent-recall -- npx -y agent-recall-mcp

For clients that take MCP JSON directly, the README supplies this block. The server name is agent-recall and the command is npx with -y and the package agent-recall-mcp:

json
{ "mcpServers": { "agent-recall": { "command": "npx", "args": ["-y", "agent-recall-mcp"] } } }

With the server registered, the first message of a new session is the loop itself. The README words it as instructions to the agent, not as CLI commands:

code
At the start of a session, call session_start to load context.
When the human corrects you, call remember with type "correction".
At the end of a session, call session_end to compound what you learned.

What you should see is a session_start injection rather than an empty context, and a stored record after the first correction. The README does not document what a fresh install prints on the first call, so treat the first session as a check that the tool names resolve and the store is writable. The repository also ships a Dockerfile that builds the MCP server workspace and runs it with node on stdio, which is the path to take if you would rather not install the workspace locally.

The published numbers are the honest limitation

The strongest reason to read this repository carefully is that it publishes numbers that argue against its own marketing. On its own live corpus, dated 2026-07-03, correction capture recall in a dual-blind audit of 59 items was 35.3 percent, with a confidence interval from 17.3 to 58.7. Correction transfer recall on an offline bench was 0 of 4, with a Wilson interval of 0 to 49 percent. The README labels the pre-reset heed rate of 92.5 percent as instrument-biased and says not to cite it; the post-reset evidence-grounded figure is 0 of 3 events, which the README calls the correct starting point rather than a regression.

The project's own explanation is density. It reports 32 active corrections across 19 projects and says that is too sparse to front-run mistakes, a data problem it says was confirmed five times by internal experiments, not a retrieval architecture problem. The README also states that transfer recall cannot support a point estimate below 39 classes, citing its benchmark spec.

That is a real constraint on adoption. If your team makes a handful of distinct corrections a month across a few repositories, you are in the same sparse regime the project measured, and the ledger will mostly record rather than prevent. AgentRecall-X is also the wrong tool if you need memory that works without an MCP-capable client, or if you want a managed service: the default configuration keeps data local, and the README does not describe a hosted offering.

How it differs from Mem0, Zep and Letta

The README names the crowded field directly: Mem0 at roughly 60K stars, Graphiti/Zep at roughly 28K, Supermemory at roughly 28K and Letta at roughly 24K. Its criticism is not that those systems retrieve badly. It is that their published numbers come from the same two or three retrieval benchmarks and are hard to reproduce independently.

The difference in approach is what gets stored and what gets measured. Mem0 and Letta style memory stores hold facts and preferences the agent chose to keep, retrieved by similarity. AgentRecall-X stores a correction with a severity, an evidence field and a proof-confidence value, and then tracks a second event: whether the same situation later produced a heeded or recurred outcome. The README describes the ledger as a governed data model, corrections-export/v1, with scrubbed egress and retraction, and says any engine can integrate against it.

So the honest comparison is not which one remembers more. It is that you can point AgentRecall-X at a number and ask what it means, because the artifacts regenerate from the repository. The README points to docs/eval/REPRODUCE.md and to scripts/eval/baselines/rmr-baseline-2026-07-03.json and scripts/eval/baselines/correction-transfer-real-2026-07-03.json. Whether you agree with the 35.3 percent capture figure, you can rerun it. That is a different posture from a leaderboard entry.

Licence, maintenance and what an upgrade actually costs

The repository is MIT licensed, which permits commercial use and modification provided the copyright notice and permission notice are preserved. That is the whole of the licence implication here; the file is LICENSE at the repository root, and anything beyond reading it is a question for your own counsel.

The project is not archived, and the last push was on 2026-09-08, which is recent enough that the codebase is moving. The release history shows a fast cadence: v3.4.48 on 2026-09-08, v3.4.47 on 2026-08-31, and v3.4.40 on 2026-07-27 with the note "Naming at scale: classifier fix, hygiene trash scan, root MANIFEST, hot-path perf". Three patch releases inside six weeks, with a minor bump that changed naming behaviour, tells you upgrades are not purely cosmetic.

The upgrade cost is structural rather than financial. The workspace layout means core, mcp-server, sdk and cli version together, and the sync-readme script copies the root README into all four packages, so the published package docs follow the root. A v3.4.x bump can therefore change both the record format and what your agent reads at session_start. The repository ships UPGRADE-v3.4.md and migration.sql at the root, which is where a schema change would be handled; the README does not document rollback, so plan a way back before you run a migration against a ledger you care about.

Editorial conclusion

Adopt AgentRecall-X if you run Claude Code or another MCP client across many sessions and you want corrections stored as structured records with a severity and an outcome, not as free-form notes. Do not adopt it if you need a retrieval benchmark winner, a hosted service, or a memory layer that works without a client that speaks MCP. Before you commit, read docs/eval/REPRODUCE.md and regenerate the numbers in UPDATE-LOG.md yourself, then check whether your own corpus is dense enough to front-run a mistake at all: the project reports 32 active corrections across 19 projects as too sparse, and that density, not the retrieval design, is what its own transfer benchmark scores against.

Frequently asked questions

How do I install AgentRecall-X for Claude Code?

The README gives a single command, claude mcp add --scope user agent-recall -- npx -y agent-recall-mcp, which registers the server at user scope. Clients that take MCP JSON directly can use the block with the command npx and the args -y and agent-recall-mcp. After that, the first message of a new session should instruct the agent to call session_start, remember with type "correction", and session_end.

Does AgentRecall-X send my data to a cloud service?

The README's badge row states cloud zero by default, and the default install runs the MCP server locally through npx. The corrections ledger is described as a governed data model with scrubbed egress, which implies export paths exist, but the README does not document a hosted backend. Anything beyond the default configuration is not described in the repository files.

How is AgentRecall-X different from other agent memory tools?

The README says other tools store facts the agent chose to keep, while AgentRecall-X stores corrections as structured records with severity, evidence and outcome tracking. It also claims no public benchmark measures whether a captured correction changes what a fresh agent does in a new session, and that it built a measurement harness for that second step. The README names Mem0, Graphiti/Zep, Supermemory and Letta as the crowded field it is distinguishing itself from.

Official sources

  1. Goldentrii/AgentRecall-X on GitHub
  2. Issues
  3. License: MIT
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/goldentrii-agentrecall-x.svg)](https://hysenlabs.com/projects/goldentrii-agentrecall-x)