Reflexio persists corrections as playbooks, and sells the rest as a service
Make your agents improve themselves. Reflexio is an AI agent self-improvement harness that enables your AI agents to continuously learn from real user interactions.
At a glance
- What is it?
- An Apache-2.0 self-improvement harness that turns user corrections and expert answers into scoped or shared playbooks without retraining, benchmarked on four of five GDPVal tasks against a warm baseline, with enhanced retrieval and continuous RL-driven improvement reserved for the paid tier.
- Who is it for?
- Reflexio suits a team already shipping an agent with real traffic, since the input it needs is conversations and corrections you are already generating, and the output is a set of playbooks rather than a retrained model you have to validate. It does not suit a prototype with no usage data, and it will not give you the enhanced retrieval or continuous RL-driven improvement that the managed tier advertises.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 5 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 26, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Conversations in, playbooks out, with humans as a second source
Reflexio is described as an agent self-improvement harness. You publish conversations from your agent into it, and it converts user corrections and interaction signals into improved decision-making, persists strategies that worked, and captures successful execution paths for reuse. The pitch is that the moat for an agent is what it learns from every interaction it handles.
There are two input streams, and the second is the interesting one. Agent conversations are the obvious source. The other is a human expert publishing ideal responses alongside what the agent actually said, from which Reflexio extracts actionable playbooks by reading the difference between them. That turns a review process you probably already do into training data, and it is the mechanism that produces a lesson before a user has to suffer the mistake.
Inside, the work splits into user profiles, playbook extraction, playbook aggregation across users, and success evaluation. The stated promise is that an agent becomes smarter, faster and more effective at domain-specific tasks as it is used, with no retraining: an approved playbook improves behaviour for everyone who loads it.
Learnings stay scoped to one user until they recur
The scoping rule is what keeps this from becoming a shared hallucination. A correction from one user becomes a learning scoped to that user. Only when the same lesson recurs across users is it aggregated into a shared playbook, and a shared playbook still has to be approved before it changes behaviour for anyone.
So there are three tiers with different blast radii: a per-user learning that affects one interaction pattern, an aggregated candidate that sits in a queue, and an approved playbook that applies to everyone. Approval is the gate, and it is a human decision, which is the correct place for it given that a bad aggregated rule would degrade every conversation rather than one.
The retrieval side runs three roles worth knowing about, generation, evaluation and a `should_run` decision, which is what decides whether a memory is relevant enough to be worth fetching at all. That third role is unusual: most memory systems decide what to store and what to return, while a separate `should_run` gate in the middle can decline to spend retrieval at all.
The benchmark runs on four of five tasks, against a warm baseline
The headline numbers are 81% fewer planning steps and 72% fewer tokens, measured on real knowledge-work tasks from OpenAI's public GDPVal benchmark, with the writeup in `benchmark/gdpval/RESULTS.md`.
Two qualifiers in the same paragraph matter more than the percentages. The results cover four of the five tasks, and the visible text does not say why the fifth is missing. And the comparison is against a warm baseline, defined as the same agent re-running the task after it has already learned from itself, which is a harder baseline than a cold agent and a more flattering one to state.
The framing is at least explicit about what that means: the savings come on top of what an already self-improving agent learns on its own. Running on four of five tasks against a warm baseline means the claim is narrower than the headline suggests, and a reader deciding whether to install this should read `RESULTS.md` before treating 81% as a planning-time saving they will see.
The paid tier advertises exactly what the open source CLI omits
There is a hosted version at reflexio.ai with a free start and a 30-day Pro trial, and the README describes its advantages as enhanced retrieval and continuous RL-driven improvement, without infrastructure to operate yourself.
That list is the boundary. Neither enhanced retrieval nor continuous RL-driven improvement is part of the open-source CLI described in the same document, which runs on SQLite with local embeddings and a local cross-encoder. So the two retrieval improvements the project considers important enough to sell are the two you do not get by installing the package, and the honest way to evaluate the open-source edition is to assume default retrieval quality.
The rest of the difference is operational rather than architectural. The open-source CLI stores data under `~/.reflexio`, runs with SQLite storage and no authentication by default, and needs exactly one LLM provider key set. The hosted version is a service with the trial attached.
Three ports, and a reranker that disappears without an error
Running the CLI locally starts several services. From PyPI it is two commands:
pip install reflexio-ai
# start/stop services. data saved under ~/.reflexio
reflexio services start # API (8061), inference (8069), SQLite storage
reflexio services stop # Stop all servicesFrom a source checkout the same launcher brings up a docs service on 8062 as well, after `uv sync` and `npm --prefix docs install`. There are also `python -m reflexio.cli services start` and a `run_services.sh` for people who would rather not go through the launcher.
The configuration detail worth planning around is how quietly retrieval degrades. When the backend is selected, a local inference service starts on 8069 unless `REFLEXIO_EMBEDDING_SERVICE_URL` points elsewhere, and it serves embeddings plus an optional cross-encoder reranker. `REFLEXIO_RERANK_ENABLED` defaults to true, and setting it to false skips reranker loading entirely. If that loopback service is unavailable, automatic reranking is skipped silently and retrieval order is preserved; an unavailable remote service is reported but search still fails open. A silent skip means ranking quality changes with no signal in your logs, which is the behaviour to alert on yourself.
Local embeddings through ONNX Runtime, with a padding shim
The default local model is `local/minilm-l6-v2`, run through ONNX Runtime, tokenizers and NumPy directly, and the README states plainly that ChromaDB is not required.
There is a compatibility detail in there that will matter if you compare vectors against anything else. Existing 384-dimensional MiniLM vectors are retained and padded to 512 dimensions, and the model cache lives under `~/.cache/chroma/onnx_models/all-MiniLM-L6-v2/onnx`, so the path still carries the old vector database's name even though the database is gone. A cold cache downloads roughly 80 MB.
Vector search itself runs through the sqlite-vec loadable extension for native KNN, with a pure-Python fallback in `sqlite_storage/_base.py` for platforms whose SQLite build cannot load extensions. That fallback is the reason the prerequisites ask for Python's linked SQLite runtime at 3.35.0 or newer rather than the standalone sqlite3 command line tool, and you can check what you actually have:
python -c "import sqlite3; print(sqlite3.sqlite_version_info)"One provider key is enough, and the priority list picks for you
The example environment file is the clearest statement of the defaults. The open-source CLI runs with SQLite storage and no auth, so the only required value is one LLM provider key, and ten are listed: OpenAI, Anthropic, OpenRouter, Gemini, MiniMax, DeepSeek, DashScope, ZAI, Moonshot and xAI.
If you set more than one, priority decides, and the order is not alphabetical or cheapest-first: Anthropic comes before Gemini, OpenRouter, DeepSeek, MiniMax, DashScope, xAI, Moonshot, ZAI and finally OpenAI. So adding a key can change which provider your agent talks to and what it costs without changing any code.
Roles resolve separately. Generation, evaluation, `should_run`, pre-retrieval and embedding each need a model, and if the top-priority provider has no embedding support the next embedding-capable one is used, currently OpenAI or Gemini. Resolution order is an org-level LLMConfig override first, then non-empty values in a checked-in JSON settings file, then the auto-detected default.
For releases, the Makefile drives two packages, the main one and a client distribution, with bump, publish, dry-run and TestPyPI targets that require uv and a publish token.
Editorial conclusion
Reflexio suits a team already shipping an agent with real traffic, since the input it needs is conversations and corrections you are already generating, and the output is a set of playbooks rather than a retrained model you have to validate. It does not suit a prototype with no usage data, and it will not give you the enhanced retrieval or continuous RL-driven improvement that the managed tier advertises. Verify first that your provider key gives you the role you need, since default provider priority picks Anthropic first and model resolution reads an org-level config, that you accept reranking being skipped without an error when the local inference service is down, and that a 384-dimensional embedding padded to 512 does not distort your similarity search.
Frequently asked questions
What does Reflexio do with user corrections?
It converts them and other interaction signals into improved decision-making so the agent stops repeating the same mistake, and it persists strategies that worked so they are reused instead of rediscovered. Successful execution paths are captured the same way.
How does Reflexio share learnings across users?
A correction from one user stays scoped to that user. Lessons that recur across users are aggregated into a candidate shared playbook, and an approved playbook changes behaviour for everyone without retraining the model.
What did the Reflexio GDPVal benchmark measure?
On four of the five real knowledge-work tasks from OpenAI's public GDPVal benchmark, a median reduction of 81% in planning steps and 72% in tokens, on a Hermes agent running MiniMax-M2.7, measured against a warm baseline where the agent had already learned from itself. The writeup is in `benchmark/gdpval/RESULTS.md`.
What do I need to run Reflexio locally?
Python 3.12 or newer, and Python's linked SQLite runtime at 3.35.0 or above. `pip install reflexio-ai` then `reflexio services start` brings up the API on 8061, inference on 8069 and SQLite storage under `~/.reflexio`. One LLM provider key is the only configuration required.
Does Reflexio need ChromaDB or a GPU?
Neither is required. The default local model runs on ONNX Runtime, tokenizers and NumPy directly, and the README states ChromaDB is not required, though the 384-dimensional vectors are padded to 512 dimensions and the cache path still sits under a chroma directory.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/reflexioai-reflexio)