# Recall: fully-local project memory for Claude Code and OpenCode

> Recall appends every Claude Code session to a local log and condenses it with a vendored TF-IDF plus TextRank summarizer, so resuming a project costs no model tokens. The trade-off is that the digest is extractive, not written by a model.

**raiyanyahya/recall** — Stop wasting tokens and re-explaining your project every session. Recall gives Claude Code , Opencode durable memory — entirely offline.

- Repository: https://github.com/raiyanyahya/recall
- Website: https://recallplugin.dev
- Stars: 752 · Forks: 44
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/raiyanyahya-recall

## The cold-start problem Recall is aimed at

Every Claude Code session starts with no memory of the previous one. The README states the problem plainly: "Claude Code starts every session cold." The practical cost is that you re-explain the project, the goal and the state of the work at the top of each session, and that explanation is billed as input tokens every time.

Recall is for people running Claude Code locally on a subscription, and the README says so directly: "It's built for people running Claude Code locally on a subscription." The design follows from that audience. If summarization were done by a model call, persistent memory would consume the same usage credits the tool is meant to protect. So the summarizer is a classical Python algorithm instead of an LLM, and the README frames the benefit as "no more re-explaining the project each session" without a metered summarizer running up a bill.

The second audience is anyone whose transcripts cannot leave the machine. Transcripts contain code, file paths and sometimes secrets. Recall's answer is that nothing is sent anywhere, and the repository carries a PRIVACY.md for the full policy. The README contrasts this with tools that "pipe your context to a model endpoint" and claims Recall "makes a privacy guarantee they can't." That is a positioning statement, not a measured result, but the architecture supports it: there is no API key to configure and no external model in the loop.

## Two files, one append-only log and one rewritten digest

Recall writes into a `.recall/` directory inside your project. Two files matter. `history.md` is the log: append-only, capturing each session as it happens, including your prompts, Claude's replies, the files touched and the commands run. `context.md` is the summary: overwritten by the local summarizer and described as the condensed "where are we right now" that you load into the next session, covering goal, summary, next steps and open threads, files touched, and where you left off.

The split matters because the two files have different lifecycles. The log grows monotonically and is never regenerated. The digest is disposable and reproducible: it is overwritten each time the summarizer runs, so a bad summary costs you nothing but a re-run. Because both are plaintext markdown under `.recall/`, they diff and can be shared like any other file in the repository.

Capture is incremental. According to the README, the `Stop` and `SessionEnd` hooks append only new activity, so a long session does not re-append turns already written. Capture happens locally, with no network step described anywhere in the flow.

## How the summarizer turns a transcript into context.md

`scripts/summarizer.py` is an extractive summarizer, and the README describes its pipeline in four steps: TF-IDF sentence vectors, a cosine-similarity graph between sentences, TextRank (PageRank power iteration) over that graph to score sentences, and finally keeping the top N in their original order.

Extractive means the summary is assembled from sentences that already exist in the transcript. Nothing is rewritten, paraphrased or inferred. If your session contains a clear statement of intent, it can surface in `context.md`; if it does not, no amount of ranking will invent one. That is the central design trade-off, and it is worth being explicit about: Recall gives you a deterministic digest, not a model's interpretation of your work.

Around the ranked sentences, `context.md` wraps deterministic facts pulled from the transcript and from git: the goal (your first ask), files touched, commands run, where you left off, and `git diff --stat`. Those fields are not scored or ranked, which is why they tend to be the most reliable part of the file. The `git diff --stat` line in particular tells you what actually changed on disk, independent of what the conversation said.

The implementation is vendored. The README states that "No installs are required" and that the whole TF-IDF plus TextRank implementation lives in `summarizer.py`. If `numpy` is importable it is used to vectorize the math and is faster on big sessions; if not, an identical pure-Python TextRank runs. The README is explicit that numpy is "an optional accelerator, never a requirement" and that the save output reports which path ran. The pyproject.toml comment agrees: "Recall ships no installable package and has no runtime dependencies."

## Installing Recall and running a first save

Recall is a Claude Code plugin, not a Python package. The pyproject.toml header says the file is "Tooling config only" and that Recall "ships no installable package," so there is no `pip install` step and no console entry point to invoke. Installation means loading the plugin into Claude Code. The README does not spell out the plugin-loading command, so check the homepage at recallplugin.dev or the `.claude-plugin/` directory in the repository for the manifest and the current install instructions.

Configuration is a file in your project root. The README calls it `recall.config.json` and the repository ships one at top level, so you can copy it rather than write it. The documented keys are `output_dir` (default `".recall"`), `capture_history` (default `true`), and `auto_save_context` (default `"off"`, with `"on_end"` as the alternative). Setting `auto_save_context` to `"on_end"` makes `context.md` regenerate every time a session ends, so you never call the save command manually. Leave it at `"off"` if you would rather decide when the digest is rebuilt.

Once the plugin is loaded, work normally for a session. The `Stop` and `SessionEnd` hooks append activity to `.recall/history.md` as you go. When you want the digest, run the save command:

```bash
/recall:save
```

That runs the local summarizer against `history.md` and overwrites `context.md`. The save output tells you whether the numpy path or the pure-Python path ran. Two other commands cover inspection: `/recall:show` prints `context.md`, and `/recall:log` tails `history.md`. At the start of the next session, the `SessionStart` hook surfaces `context.md` and has Claude ask you two things: whether to resume from the saved context, and whether to keep logging this session.

## Where Recall sits next to CLAUDE.md and --resume

Claude Code already has memory features, and Recall does not replace them. The README is careful about the distinction. `CLAUDE.md` and the `#` shortcut are hand-written memory: rules and notes you curate, loaded as instructions Claude follows. They describe how you want Claude to work, but they require manual upkeep and they do not record what happened.

`--continue` and `--resume` take the opposite approach. They replay a prior conversation at full fidelity, which means reloading the whole transcript and paying for it in tokens, and the state is tied to your local session history on one machine rather than a portable digest. Context compaction sits inside a single session and is not a durable record you reopen days later.

Recall occupies the gap: an automatic, deterministic record of what each session did, condensed into a compact resume point. The README puts the resume cost at roughly 1 to 2K tokens for the compact digest, against a full transcript replay for `--resume`. Treat that figure as the project's own characterization rather than a measured benchmark.

There is one more difference worth naming. The README says Claude treats `context.md` as "untrusted reference data" rather than as instructions. That fencing is a sensible default, since a transcript can contain text that was never meant as a directive, but it also means the digest informs the session rather than steering it. If you want persistent rules, they still belong in `CLAUDE.md`.

## The extractive summary is the limitation

The clearest failure mode follows from the algorithm. TextRank scores sentences by centrality in a similarity graph, so it favors sentences that resemble many other sentences. In a session where the important decision is stated once, briefly, and never repeated, that sentence can lose to verbose boilerplate that shares vocabulary with everything around it. The deterministic fields (goal, files touched, commands run, `git diff --stat`) are the part of `context.md` least exposed to this problem.

Recall is also the wrong tool if you want a summary that reasons. An LLM-written digest can say "the migration stalled because the schema change was reverted" when no sentence in the transcript says that. Recall cannot, by construction. It selects, it does not synthesize.

There are practical boundaries too. The README does not document rollback for `context.md`, though the file is overwritten on each save and the log is append-only, so the prior digest is not retained by the tool itself. Nothing in the README describes multi-machine sync: `.recall/` is plaintext in your project, so sharing it is a git problem, not a Recall feature. OpenCode support arrived in v0.4.0 and is labeled "opt-in" in the release title and the README badge, so it is not the default path. Finally, the project's own test configuration excludes `scripts/session_start.py` from coverage because it is "a thin lifecycle entry point exercised only inside a live Claude Code session." That is an honest admission that the session-start path cannot be verified by the unit suite.

## Alternatives and the difference in approach

The most direct alternative is doing nothing beyond the built-ins and relying on `--continue` or `--resume`. The difference is fidelity against cost: `--resume` gives you the full prior conversation, which is the highest-fidelity option available and needs no summarizer at all, but it reloads the whole transcript and the state lives on one machine in local session history. Recall trades that fidelity for a compact, portable, plaintext digest costing roughly 1 to 2K tokens to load.

A second alternative is any memory tool that pipes context to a model endpoint for summarization. The difference is where the compute happens. Those tools can produce a written, reasoned summary because a model does the writing. Recall cannot, and in exchange your code, paths and whatever else is in the transcript never leave the machine. For a repository containing credentials or client code, that exchange is often the deciding factor.

A third option is `CLAUDE.md` alone. It is free, it loads as instructions rather than reference data, and it is fully under your control. What it does not do is record what happened in a session, which is precisely the gap Recall targets. The README's own framing is that `CLAUDE.md` is how I want you to work, and Recall is here is what we did last time and where we stopped.

## Maintenance, upgrade cost and the MIT licence

The last push to the repository was on 2026-09-07, roughly two weeks before this writing, and the repository is not archived. Releases have been steady through the middle of 2026: v0.3.5 on 2026-06-22, v0.3.6 on 2026-06-25, and v0.4.0 on 2026-07-18, which added opt-in OpenCode support.

Upgrade cost is unusually low, and that follows from the packaging decision. Because Recall ships no installable package and has no runtime dependencies, there is no dependency tree to reconcile and no version pin to bump. Upgrading means updating the plugin, not resolving a lockfile. The vendored summarizer means a numpy upgrade cannot break the pure-Python path, and the pyproject.toml comment states numpy is "never required."

The project also treats hook failure as a design constraint rather than an edge case. The bandit configuration in pyproject.toml records that the lifecycle hooks deliberately `except Exception: pass` so that "a failure in Recall can never crash a Claude Code session," and notes this is "a hard design requirement" documented in CONTRIBUTING.md. The same file documents that the only subprocesses are git and the opencode CLI, invoked with a static argv list rather than `shell=True`, with git additionally hardened against untrusted-repo configuration in `scripts/common.py`.

Recall is MIT licensed, which permits commercial use, modification and redistribution provided the copyright notice and licence text are retained. That is a summary of the licence, not legal advice; read LICENSE and PRIVACY.md before shipping it inside a regulated workflow.

## Conclusion

Recall fits people running Claude Code or OpenCode on a subscription who want a resume point without paying a summarizer per session, and who accept an extractive digest over a model-written one. Skip it if you need summaries that reason about your code rather than rank its sentences, or if you already get what you need from CLAUDE.md plus --continue. Before adopting, run /recall:save once and read the generated .recall/context.md to see whether the ranked sentences are actually the ones you would have written.

## FAQ

### How do I install Recall for Claude Code?

Recall is a Claude Code plugin rather than a Python package, and the pyproject.toml states it ships no installable package, so there is no pip step. The README does not give the plugin-loading command itself; check recallplugin.dev or the .claude-plugin/ directory for the manifest and current instructions.

### Does Recall send my transcripts to an API?

No. The README states that nothing leaves your machine, that there is no API key and no external model, and that summarization is done by a classical Python summarizer running locally. PRIVACY.md holds the full policy.

### What is the difference between history.md and context.md in Recall?

history.md is the append-only log that the Stop and SessionEnd hooks write to as a session happens, covering prompts, replies, files touched and commands run. context.md is the digest that the local summarizer overwrites, holding the goal, summary, next steps, files touched and where you left off.

### Does Recall need numpy or any other dependency?

No. The README says the whole TF-IDF plus TextRank implementation is vendored in scripts/summarizer.py, and numpy is described as an optional accelerator that is never a requirement. If numpy is importable the save output reports that path; otherwise an identical pure-Python TextRank runs.

## Sources

- [License: MIT](https://github.com/raiyanyahya/recall/blob/master/LICENSE)
- [Project website](https://recallplugin.dev)
- [raiyanyahya/recall on GitHub](https://github.com/raiyanyahya/recall)
- [README](https://github.com/raiyanyahya/recall/blob/master/README.md)
- [Releases](https://github.com/raiyanyahya/recall/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/raiyanyahya-recall
