Model or dataset
BayramAnnakov/claude-reflect avatar
BayramAnnakov/claude-reflect

claude-reflect: a memory system for a coding agent

A self-learning system for Claude Code that captures corrections, positive feedback, and preferences — then syncs them to CLAUDE.md and AGENTS.md.

1,703 stars150 forksPythonMIT

At a glance

What is it?
The interesting design here is not that the tool stores what you tell it, it is that the automated half only ever writes to a queue. Hooks fire on every prompt and after every git commit, regexes catch corrections and praise in real time, and a model pass cleans them up only when you type `/reflect`. What lands in your instruction files is something a human pressed Apply on.
Who is it for?
Adopt claude-reflect if you correct the same assistant often enough that repeating yourself is costing you, and run `/reflect --targets` first to see the write surface, because a learning accepted into `~/.claude/CLAUDE.md` changes the assistant's behaviour in every project you open. Do not adopt it expecting dependable capture: the real-time stage is pattern matching, and the release title for v3.2.0 is two silent failures fixed, which is the shape this failure mode takes.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 12 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The automated half only writes to a queue

The pipeline diagram in this README is three boxes and the annotations under them carry the whole design. You correct Claude Code, a hook captures it to a queue, and `/reflect` adds it to CLAUDE.md. Under those three boxes: automatic, automatic, manual review.

So there are two stages and only one of them is unattended. The capture side runs on every prompt you send, watching for the shape of a correction. Nothing it finds reaches your instruction files. The processing side is something you invoke, and when you do, the tool shows you a table of what it found with three choices per row: apply it, edit the text before applying, or skip it.

That is the property that makes a tool like this usable at all. A system that watches every prompt you write is going to catch things you did not mean for it to catch, and the only question is what happens when it is wrong. Here the answer is bounded: nothing is written, and there is a command that throws the queue away.

The commands list is worth reading as a safety inventory rather than a feature list. There is `/reflect --dry-run` to preview without applying, `/reflect --targets` to list the config files that would be updated, `/view-queue` to see what is pending, `/skip-reflect` to discard the lot, and `/reflect --review` which shows the queue with confidence scores and a decay status. Five ways to look before you write, for a tool whose output is a set of instructions an assistant will follow for ever.

The remaining flags are narrower. `--scan-history` goes back over past sessions for things the real-time hooks missed. `--dedupe` finds and consolidates similar entries already in your CLAUDE.md, which is the housekeeping you need once the file has been growing for months. `--include-tool-errors` widens the scan to tool execution failures, not just things a human typed.

Regex on every prompt, a model pass when you ask

Detection is described as a hybrid approach, and the two halves exist for reasons that are visible once you put them next to each other.

The first half is regex, running in real time during a session. It looks for corrections such as "no, use X", "don't use Y", "actually..." and "that's wrong"; positive feedback such as "Perfect!", "Exactly right" and "Great approach"; and explicit markers like "remember:", which carry the highest confidence. This has to run on every prompt, so it has to be fast and it has to work offline.

The second half is a semantic pass, and it runs during `/reflect` rather than during the session, because it costs a model call. Its stated jobs are three: understanding corrections in any language, filtering out the false positives the regexes produce, and extracting a concise actionable statement from whatever was said. The example given is a Spanish correction, "no, usa Python", which the regexes cannot match and the semantic pass catches.

That division is the right one. Anything running on every prompt has to be cheap enough not to be noticed and deterministic enough to be predictable, which rules out a model. Anything filtering a queue you chose to review is exactly where a model call belongs.

Two details are worth sitting with. First, praise is captured alongside correction, so the memory is not only a record of what went wrong. The key features table calls this skill improvement, where a correction during a deploy changes the deploy command itself. A memory that records only failures teaches an agent what not to do; one that records what worked teaches it what to repeat.

Second, confidence runs from 0.60 to 0.95 and the final score is the higher of the regex and semantic confidence. Taking the maximum is a choice with a direction: when the two disagree, the more confident one wins, so a confident pattern match can outrank a semantic judgement. That biases the queue towards surfacing something for you to look at, which is the right way round for a tool where the human is the final filter, but it does mean the review queue will be longer than a minimum-vote rule would give you.

What a learning is allowed to write to

This is the part to read before installing the plugin, because it defines the blast radius.

An approved learning is synced to several places. There is a global file in your home directory that applies to all projects. There is a project file. There are files in subdirectories, auto-discovered by a glob. There are the command files in the project's commands directory, so a correction can rewrite an existing slash command rather than only adding a note. There is an `AGENTS.md` if one exists, and the README notes that file works with Codex, Cursor, Aider, Jules, Zed and Factory, so a correction accepted once propagates to whichever agent you happen to be using. And then the recent addition: any document reachable from those files through a filename include or a markdown link, at bounded depth, with cycles handled and code blocks and external URLs skipped.

That last item is the one the release history explains. Version 3.2.0 is titled two silent failures fixed, referenced docs as targets, which tells you the feature exists because a user had a CLAUDE.md that delegated its guidance to a standards file, and the learning landed in the wrong file or nowhere. It is a sensible addition for the same reason, and it is also an amplification of the blast radius: the write surface is now a graph rather than a list, and `/reflect --targets` exists to make that graph inspectable before you commit to it.

The concrete risk is the global file. A learning you apply to your home-directory config changes the behaviour of an assistant in every project you open from then on, including the projects where the correction was about something else. Nothing here prevents that, which is why `--targets` and `--dry-run` are worth running on the first invocation rather than the twentieth.

Four hooks, two of which exist to stop you losing data

The capture stage is four Python scripts, and the trigger for each one tells you what problem it solves.

A session-start hook shows a reminder about pending learnings, so a queue that built up while you were not looking does not go unnoticed. A capture hook runs on every prompt and does the detection work described above. A hook that fires before compaction backs up the queue and tells you about it, and that is the most interesting of the four, because compaction is the point at which in-flight session content is discarded. If the queue lived only in the session transcript, compaction would be a data-loss event, so the hook writes it out first.

The fourth fires after a git commit and reminds you to run the processing step once you have finished a piece of work. That is a workflow decision rather than a technical one: a commit is a natural boundary, and processing learnings at a boundary is easier to remember than processing them whenever you happen to think of it.

The prerequisites are minimal by design, which is worth knowing. The Claude Code CLI, and Python 3.6 or newer, which the README notes is included on most systems. Platform support is stated as macOS and Linux fully supported, and Windows fully supported with native Python and no WSL required. The repository carries its own CLAUDE.md and SKILL.md at the root, so the project is configured to be worked on by the tool it provides.

The installation path is the plugin marketplace route, and it ends with a restart:

bash
# Add the marketplace
claude plugin marketplace add bayramannakov/claude-reflect

# Install the plugin
claude plugin install claude-reflect@claude-reflect-marketplace

# IMPORTANT: Restart Claude Code to activate the plugin

The restart is not cosmetic, because the hooks have to be installed into the assistant's own configuration before they can fire. Your first `/reflect` then prompts you to scan past sessions, which is the `--scan-history` path and the fastest way to see whether the tool has anything worth keeping before you commit to reviewing it routinely.

Pattern mining, and where the thresholds are yours

The second command is `/reflect-skills`, and it does something different: it reads your session history looking for repeated requests, and proposes turning them into reusable commands.

The example in the README is a request made twelve times, which then becomes a proposal for a daily review command. The reasoning is sound and unglamorous: a request you have typed a dozen times is a command you should have written, and counting repetitions is a cheap way to find it.

The controls are what matter. The default window is the last fourteen days, and `--days N` changes it. `--project <path>` scopes the scan to one project. `--dry-run` previews the patterns without generating any files. And `--all-projects` searches across all of them for patterns that repeat across your whole working life.

That last flag is where the tool stops being conservative. A pattern that appears in one project is likely a fact about that project. A pattern that appears across every project you have touched in a window is more likely to be an artefact of how you work rather than a task worth codifying, and a proposal generated from it becomes a command in a directory that applies everywhere. Nothing stops it, and the review step is a proposal rather than an application, so the human still decides. But the failure mode is a global command built from a coincidence, and `--project` is the setting to leave on until you have seen a good cross-project proposal.

The same caution applies to the corrections path. The features table lists multi-language understanding as a capability, which means a correction made in any language can become a permanent instruction. That is genuinely useful and it widens the surface of things that can end up in a global file without you noticing the language they were written in.

Version cadence, and where this should not be installed

The licence is MIT, the last push was on 2026-09-20, and the repository is not archived.

The release history has a shape worth noting. Version 2.6.0 and version 3.0.0-rc.1 were both released on 2026-02-13, hours apart, the first titled with a session retention warning and the second with full memory hierarchy integration. Then nothing until 3.2.0 on 2026-09-20. So a release candidate was published to the release list as a release, and there is a seven-month gap between it and the next version, which skips 3.0.0 and 3.1.0 on the way to 3.2.0. For a tool whose whole output is a set of instructions an assistant follows, that cadence is the thing to factor into an evaluation.

The tests badge claims 322 passing, which is the project's own claim and a reasonable one to weigh alongside the fact that a test suite does not tell you whether the regexes match your vocabulary.

Where not to install it is straightforward enough. Not on a machine whose Claude configuration you cannot afford to have rewritten, since Apply writes to your home directory and can follow links into files you did not know were part of the graph. Not in a shared or work account until you have read `--targets` output on it. And not as a substitute for writing down the conventions that matter, because the mechanism here is pattern matching on prompts plus a model pass, which will capture the corrections you make and not the ones you never think to make.

Editorial conclusion

Adopt claude-reflect if you correct the same assistant often enough that repeating yourself is costing you, and run `/reflect --targets` first to see the write surface, because a learning accepted into `~/.claude/CLAUDE.md` changes the assistant's behaviour in every project you open. Do not adopt it expecting dependable capture: the real-time stage is pattern matching, and the release title for v3.2.0 is two silent failures fixed, which is the shape this failure mode takes. Read the Python 3.6 floor before you package it, and prefer `--project` over `--all-projects` until you trust the pattern mining.

Frequently asked questions

What does claude-reflect do in Claude Code?

It captures the corrections you make to the assistant, along with positive feedback and explicit markers, queues them automatically, and syncs what you approve into instruction files. Hooks run on every prompt and after each git commit, and the queue is processed only when you run the `/reflect` command.

How does claude-reflect decide whether something is a correction?

With two stages. Regex patterns run in real time during a session and match phrases such as "no, use X", "actually..." and "remember:", as well as praise like "Exactly right". When you run `/reflect`, a semantic pass handles multi-language input and filters the false positives, producing a confidence score between 0.60 and 0.95.

Which files can claude-reflect write to?

The global `~/.claude/CLAUDE.md`, the project `CLAUDE.md`, CLAUDE.md files in subdirectories, the project's command files under `.claude/commands/`, an `AGENTS.md` if one exists, and any documents reachable from those files through filename includes or markdown links at bounded depth, with cycles handled and code blocks and external links skipped.

How do I install the claude-reflect plugin?

Add the marketplace with `claude plugin marketplace add bayramannakov/claude-reflect`, install with `claude plugin install claude-reflect@claude-reflect-marketplace`, then restart Claude Code so the hooks activate. You need the Claude Code CLI and Python 3.6 or newer, and Windows is supported with native Python rather than WSL.

Can I preview what claude-reflect would change?

Yes. `/reflect --targets` lists the config files that would be updated, `/reflect --dry-run` previews without applying, `/view-queue` shows what is pending, `/reflect --review` shows the queue with confidence scores and decay status, and `/skip-reflect` discards everything queued.

What does the claude-reflect skill discovery command do?

It reads your session history looking for repeated requests and proposes turning them into reusable commands, for example a request made twelve times becoming a daily review command. It defaults to the last 14 days and accepts `--days N`, `--project <path>`, `--all-projects` and `--dry-run`.

Official sources

  1. BayramAnnakov/claude-reflect on GitHub
  2. Issues
  3. License: MIT
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/bayramannakov-claude-reflect.svg)](https://hysenlabs.com/projects/bayramannakov-claude-reflect)