coleam00/claude-memory-compiler: a markdown knowledge base built from Claude Code sessions
Give Claude Code a memory that evolves with your codebase. Hooks automatically capture sessions, the Claude Agent SDK extracts key decisions and lessons, and an LLM compiler organizes everything into structured, cross-referenced knowledge articles - inspired by Karpathy's LLM Knowledge Base architecture.
At a glance
- What is it?
- The project captures Claude Code transcripts through hooks, extracts decisions and lessons with the Claude Agent SDK, and compiles them into cross-referenced markdown articles. It is designed for one person's machine, not a team wiki.
- Who is it for?
- Adopt it if you work in Claude Code daily on one machine and want your own decisions and gotchas to survive across sessions; the hooks, `flush.py` and `compile.py` are the whole system. Do not adopt it if you need a shared team knowledge base, a hosted service, or a licence you can point at, because the repository does not state one.
- Can I use it commercially?
- Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
- Is it still maintained?
- Yes. The repository last received commits 176 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem: Claude Code forgets what you decided yesterday
Claude Code sessions end and the reasoning inside them goes with them. You explain an architecture choice on Monday, and on Thursday the agent proposes the opposite. The usual fixes are manual notes, a running scratch file, or a chat history you never reopen.
claude-memory-compiler targets that gap specifically. The README frames the raw material as "your own conversations with Claude Code" rather than clipped web articles, which is the difference from the Karpathy gist it adapts. It is built for a single developer working on one codebase, someone who wants the decisions, lessons, patterns and gotchas from their own sessions to come back in a structured form. It is not a team knowledge base and the README does not present it as one.
Hooks in, markdown out: the capture and compile pipeline
The README gives the data flow as a single line, and it is worth reading literally:
Conversation -> SessionEnd/PreCompact hooks -> flush.py extracts knowledge
-> daily/YYYY-MM-DD.md -> compile.py -> knowledge/concepts/, connections/, qa/
-> SessionStart hook injects index into next session -> cycle repeatsTwo hook points matter. SessionEnd fires when a session closes; PreCompact fires before Claude Code compacts a long conversation, so a session that never cleanly ends still gets captured. `flush.py` receives the transcript and calls the Claude Agent SDK to decide what is worth keeping, then appends the result to a daily log at `daily/YYYY-MM-DD.md`. After 6 PM local time, a flush also triggers end-of-day compilation, so the day's logs become articles without a manual step.
`compile.py` turns those daily logs into concept articles with cross-references under `knowledge/concepts/`, `connections/` and `qa/`. Retrieval is handled by `query.py`, which the README describes as index-guided: the agent reads a structured `index.md` instead of doing vector search. The README states the reasoning plainly, citing Karpathy's point that at personal scale (50-500 articles) an LLM reading a structured index beats cosine similarity, and that RAG only becomes necessary past roughly 2,000 articles when the index outgrows the context window. That is a real architectural bet, and it is the part most worth disagreeing with if your corpus grows.
Installing it and getting the first daily log
There is no package to install from a registry. The README's quick start is written as an instruction you hand to your AI coding agent: clone the repository into your project, set up the hooks, and read AGENTS.md for the technical reference. The repository layout supports that, with `.claude/`, `hooks/`, `scripts/`, `pyproject.toml` and `AGENTS.md` at the top level.
If you prefer to do it yourself, the project is a Python package managed by uv, and the README says the agent runs `uv sync` to install dependencies. The declared dependencies in `pyproject.toml` are `claude-agent-sdk>=0.1.29`, `python-dotenv>=1.0.0` and `tzdata>=2024.1`, with `requires-python = ">=3.12"`. Note the tzdata dependency: the 6 PM compile trigger is local-time based, so timezone data is not decorative.
uv syncNext, the hooks. The README says to copy `.claude/settings.json` into your project, or merge the hooks into settings you already have. Merging is the safer path if you have your own hooks, and the README does not describe a merge tool. The README states the hooks activate the next time you open Claude Code. From there, conversations accumulate on their own.
To compile manually instead of waiting for the post-6 PM trigger, the README gives this command:
uv run python scripts/compile.pyTo ask the knowledge base a question, and optionally write the answer back as a new article:
uv run python scripts/query.py "question"
uv run python scripts/query.py "question" --file-backWhat you should see after a session or two is a dated file under `daily/`, and after compilation, articles under `knowledge/`.
The retrieval bet and where it breaks
Skipping embeddings is the most consequential choice here. It removes an embedding model, a vector store and the chunking decisions that come with them. The cost is that every query depends on the index fitting in context and on the agent choosing the right articles to open. The README names the ceiling itself: around 2,000 articles, when the index exceeds the context window. Below that, index-guided reading is plausible. Above it, the project offers no migration path in the README.
The second limitation is scope. Everything runs from hooks on one machine, writing markdown into one project. There is no server, no sync, and the README does not describe sharing or merging knowledge bases across people. If two engineers both run it, they get two independent stores.
The third is cost accounting. The README states that Anthropic has clarified personal use of the Claude Agent SDK is covered under an existing Claude subscription (Max, Team, or Enterprise) and that no separate API credits are needed, contrasting this with OpenClaw, which it says requires API billing for its memory flush. That is a claim about billing policy, not about compute: extraction and compilation still consume usage against your plan, and the README does not quantify how much. Verify the current policy yourself before assuming a heavy session history is free.
lint.py and the maintenance you actually signed up for
A knowledge base that only grows becomes noise, and the project acknowledges this with `lint.py`, which the README says runs 7 health checks: broken links, orphans, contradictions and staleness among them. There is a free mode for structural checks only:
uv run python scripts/lint.py
uv run python scripts/lint.py --structural-onlyThe split is telling. Structural checks (broken links, orphans) are deterministic and cost nothing. Contradiction and staleness checks require the model to read the articles, which consumes usage. The README does not state how the contradiction check decides that two articles conflict, and that is the check most likely to produce noise on a personal corpus where your opinions legitimately changed over time.
Upgrade cost is low in the ordinary sense: there are no recent releases listed, the version in `pyproject.toml` is 0.1.0, and the main moving part is the `claude-agent-sdk` dependency, which is pinned to a minimum rather than a range ceiling. The last push to the repository was on 2026-04-06, so treat the code as a snapshot rather than something receiving steady changes; if the Agent SDK changes its interface, nothing in the repository guarantees a fix. Your real maintenance burden is curating the articles and running lint, not pulling updates.
Licence and the question the repository does not answer
The repository does not state a licence. There is no LICENSE file in the top-level entries (`.claude/`, `.gitignore`, `AGENTS.md`, `README.md`, `hooks/`, `pyproject.toml`, `scripts/`, `uv.lock`), and no licence identifier appears in `pyproject.toml`. Under default copyright, that means you have no granted right to redistribute or reuse the code, whatever GitHub's interface may suggest. For personal local use on your own machine this is unlikely to matter in practice. For vendoring it into a company repository, shipping it inside a product, or building on top of it, the absence of a licence is a real blocker and you should ask the author rather than assume. I am not giving legal advice; the point is that the file is missing.
How it differs from a hosted memory feature
The obvious alternative is Claude's own memory, which the search questions suggest people are already asking about. The difference in approach is where the knowledge lives and what shape it takes. A built-in memory feature is a general store of facts about you, managed by the vendor, and you get whatever interface it offers. claude-memory-compiler produces files you own, in your repository, organized by concept with explicit cross-references and a query script you can read. You can open `knowledge/concepts/` in an editor, diff it, or delete an article you disagree with. You cannot do that with a hosted memory store.
The trade-off runs the other way too. Hosted memory works across every surface you use Claude on, with no hooks, no Python environment and no uv. This project only sees Claude Code sessions in the project where the hooks are installed. If your thinking happens in a browser chat window, none of it gets captured.
Editorial conclusion
Adopt it if you work in Claude Code daily on one machine and want your own decisions and gotchas to survive across sessions; the hooks, `flush.py` and `compile.py` are the whole system. Do not adopt it if you need a shared team knowledge base, a hosted service, or a licence you can point at, because the repository does not state one. Verify first that your Claude plan covers Agent SDK personal use, that `uv` is available, and that you are willing to merge `.claude/settings.json` hooks rather than replace your existing ones.
Frequently asked questions
Does claude-memory-compiler need API credits, or does it run on a Claude subscription?
The README states that Anthropic has clarified personal use of the Claude Agent SDK is covered under an existing Claude subscription (Max, Team, or Enterprise) and that no separate API credits are needed. It contrasts this with OpenClaw, which it says requires API billing for its memory flush. The README does not quantify how much usage extraction and compilation consume.
Where does claude-memory-compiler store the knowledge it captures?
Captured sessions are appended to daily logs at `daily/YYYY-MM-DD.md`, and compilation turns those into articles under `knowledge/concepts/`, `connections/` and `qa/`. Retrieval reads a structured `index.md` rather than a vector database, and a SessionStart hook injects that index into the next session.
How do I set up claude-memory-compiler?
The README's quick start is written as an instruction to your AI coding agent: clone the repository into your project, run `uv sync`, copy `.claude/settings.json` into your project or merge the hooks into your existing settings, and read AGENTS.md. The README states the hooks activate the next time you open Claude Code.
Can I transfer a knowledge base built by claude-memory-compiler to another machine or person?
The README does not describe any transfer, sync or export mechanism. Everything is markdown written into one project by hooks on one machine, and the README does not present the project as a shared or multi-user store.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/coleam00-claude-memory-compiler)