CodeStable: a set of skill contracts whose main selling point is what it refuses to be
CodeStable 是一套面向严肃软件工程、坚持人在环的 AI 编码工作流。它不以编排 Agent 为中心,而是组织需求、架构、特性、问题与历史决策,让 Codex 和 Claude 驱动的开发过程可控、可追溯、可持续演进。
At a glance
- What is it?
- This is a human-in-the-loop workflow for AI-assisted development that deliberately does not orchestrate agents, does not run every task through one pipeline, and does not create a second documentation system. Its substantive ideas are a taxonomy of what counts as evidence for each kind of task, and a hard cap of twenty-five items on the context the agent reads in every session.
- Who is it for?
- CodeStable is worth reading if you are running AI-assisted development on something that will still exist in two years and are frustrated by the pile of plan files, status files and summary files that agent workflows leave behind, because its position is that the diff and the test output are the evidence and everything else is optional.
- Can I use it commercially?
- Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
- Is it still maintained?
- Yes. The repository last received commits 47 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The readme spends more space on what this is not than on what it is
The positioning statement in the second paragraph is unusually blunt. This is a set of lightweight skill contracts for serious software development. It does not orchestrate a team of agents, and it does not build a second documentation system for your project.
That claim is then defended for the rest of the document, and the defence is the most interesting part of the readme. There is a section of suitability that lists four situations where this fits and four where it does not, and the second list is the one worth reading. It is not a multi-agent orchestration platform or an automatic relay system. It is not a process engine that forces every task through the same pipeline. It is not a second system standing in place of the documents, decision records, issues and pull requests you already have. And it is not a dependency you need for a one-off prototype where long-term maintenance does not matter.
Each of those four is a real product category with real users, and each is a category this project could have entered by adding features. A workflow framework that orchestrated sub-agents would be more impressive. A pipeline engine would be more useful to a team that wants consistency at the cost of flexibility. A parallel documentation store would be easier to build than a discipline about not building one. Declining all four is a coherent position: the observation is that agent workflows tend to fail not because the agent cannot write code but because nobody can tell afterwards what was decided, why, and on what evidence.
The two companion projects it recommends follow the same logic. One turns agent sessions into read-only links so a team can read what happened and hand work over. The other creates and manages sub-agents and can supply an independent reviewer. The readme is explicit that the three are complementary and that neither of the other two replaces the engineering contract. That is a market map, not a feature list, and it is the clearest indication that the author knows where their project stops.
The behaviour of the main entry point follows the same restraint. Asked to do something, it either starts work or discusses it, and the readme spells out the conditions for each. When the requirements are clear it proceeds. When it hits a specific risk it adds only the confirmation, test or review that corresponds to that risk, and explicitly does not switch on the whole process. When you want to discuss first it aligns on goals, terminology and boundaries, and it will not touch code without explicit permission.
Evidence before conclusions, with a different standard for each kind of task
The second of the three core principles is the one that does the most work, and its content is a small taxonomy.
The principle is that evidence comes before conclusions. Implementation of a feature needs design and validation proportionate to the risk. Fixing a bug goes red then green. Refactoring first establishes evidence that behaviour is equivalent. Independent review is created by the outer process as a read-only reviewer rather than by the person doing the work.
Each clause is a different gate, and the differences matter more than the shared headline. A red-to-green requirement for a bug fix is the strictest of the three, because it is the only one that demands you demonstrate the failure before you demonstrate the repair. A fix with no failing test first has not established that anything was broken, so it cannot establish that anything is now fixed. That is a checkable rule and a small number of agent workflows enforce it.
The refactoring clause is the most interesting, because behaviour equivalence is the hardest of the three to demonstrate and the easiest to assert. A refactor that passes the same test suite is not behaviour-equivalent if the tests were weak to begin with, and the principle is careful to say the evidence is established first, before the change, rather than after. Establishing it first is also what makes a refactor reviewable, because the reviewer can compare the two sets rather than take the claim on trust.
The feature clause is deliberately vaguer, and that is correct. The amount of design work and validation a feature deserves depends on how much can go wrong, so a fixed gate would be wrong in both directions. The readme is explicit that the risk determines the response, and it says this twice: once as a principle and once in the description of the entry point, which only escalates when it meets a specific risk rather than escalating by default.
The last clause closes the review gap. The reviewer is read-only, it runs in a single round, and it is created from outside the task rather than invoked by the thing being reviewed. That is three separate constraints, and each one removes a way for a review to become theatre. Read-only means it cannot fix what it finds and call it done. A single round means it cannot negotiate with the author. Created from outside means the author did not choose their own reviewer.
And then there is the statement about where people enter the process: at changes to the product contract, at significant risk, and at overall acceptance, and not at every mechanical step. A human who confirms every step has not reviewed anything, and a human who reviews nothing has not accepted anything. Three named points is a defensible number.
Twenty-five items, and the least interesting design decision in the project
Onboarding creates a directory with three things in it:
.codestable/
├── attention.md
├── lessons/
└── work/The specification for the first of them is the most concrete and most borrowable idea in this repository.
That first file holds the small number of project facts needed in every session, and the specification is a maximum of twenty-five items. Not twenty-five lines, not a page. Twenty-five items, and the readme states the cap rather than implying it.
A cap is the whole design. Anything an agent reads in every session is paid for on every session, and anything paid for on every session competes with the task for attention. A project fact file that grows without limit is the single most common way a memory system degrades: it starts as five decisions, becomes forty, and by the time it is forty the model is reading a list of things that mostly do not apply to the thing in front of it. The cost is not tokens, which is the argument people make. The cost is that the five things that do apply are now four percent of the input and get treated accordingly.
Twenty-five is small enough to be a genuine constraint. It is small enough that a team has to argue about what belongs, and an argument about what belongs in always-loaded context is a useful argument to have once a year.
The second directory holds lessons, one per file, and here the design is about lifecycle rather than volume. Each lesson moves through three states, and the states are named: observed, validated, retired. Observed means somebody saw it happen. Validated means a later session confirmed it. Retired means it stopped being true or stopped being relevant. Before a lesson is written, the existing ones are checked and duplicates merged, which is a small process that prevents the same lesson accumulating in four slightly different wordings.
The third directory is the narrowest. It exists for tasks that cross sessions, for handovers between people, and for persistent records that were explicitly asked for, and the readme is firm that it is not for anything else. In particular, discussion inside a session does not go in there.
The cap and the one-file rule together are a coherent answer to the question every memory system has to answer: what do you read, and how do you know it is still true. Twenty-five answers the first. Three lifecycle states and dedup-before-write answer the second.
Put mechanical mistakes in a test, not in a lesson
The paragraph about how new lessons get created is the part of this readme that most resembles hard-won experience, and it contains a rule that is better than the rest of the system put together.
The rule is: if an error can be mechanised, it goes into a test or a checker first. A new lesson still requires explicit authorisation, and a later session verifies it before it can be validated or retired.
So the project distinguishes two kinds of knowledge explicitly. Some knowledge is a fact about the code, and a fact about the code that can be checked should be checked automatically rather than written down for a model to remember. The rest is judgement, and judgement has to live in prose because there is nothing to assert it with. A lesson file that says do not do this thing is a request to a future reader. An assertion that fails when you do it is a guarantee.
Preferring the guarantee is obvious once stated and rarely done, because writing a test is work and writing a sentence is not. The result in most projects is a memory system full of lessons that are really just missing tests, accumulating at the rate somebody had a bad afternoon, and a codebase that still has the same bug because the lesson is read by a model that will occasionally ignore it. A project that routes mechanical errors into assertions and reserves prose for judgement will have a much smaller memory directory, and it will be a much more accurate one.
The rest of the paragraph describes the discovery process, and it is unusually restrained. The system watches for moments during a task where something crystallises, does so silently, and at an ordinary wrap-up offers at most one candidate, with evidence attached. One candidate, not five. The restraint is the point: a mechanism that surfaces three candidates every session trains people to dismiss all three, and then the mechanism is decoration.
The lifecycle is also asymmetric in a useful way. Observation is automatic, validation is not, and retirement is not. The cheap operation is the one that needs no permission. The two operations that change what the system believes require a human, which is the correct place to put the cost.
Ordinary tasks produce no documents at all
There is a sentence in the memory section that is worth quoting in full, because it is a direct rejection of the most common artefact in agent-assisted development.
Ordinary tasks do not generate CodeStable stage documents. The diff, the test output and the delivery notes are the evidence.
That is a position, not an omission. The prevailing pattern in this genre is that a run produces a plan file, then a status file, then a summary, then a changelog entry, then a retrospective, and six months later nobody can tell which of them is current. Each artefact is cheap to write and expensive to trust, because a plan written before the work and a plan updated after the work are different documents with the same name, and readers assume the second.
This project closes that door. A normal feature or bug fix leaves three things behind: the code change itself, the output of whatever verification ran, and a description of what was delivered. Those three are checkable. A plan file is not checkable, because you cannot diff a plan against reality without redoing the work.
The same logic governs the session discussion. Talk inside a session does not enter the working directory at all. Only a conclusion that is stable and worth reusing graduates, and it graduates by being written into the destination that already owns that kind of fact, which is the third principle rather than a new store.
That third principle is the one that makes the whole thing sustainable. Existing project documents, decision records, code and domain documentation keep ownership of the facts they already own. The tool adds a small number of things: a capped fact file, a lesson per file, and a cursor for active work. It does not copy anything into a parallel archive. When a stable conclusion has no obvious home, the tool asks a person to choose one rather than inventing a location.
The epic workflow extends the same idea with a two-layer structure. A permanent document holds goals, scope, acceptance, approved sub-items, key decisions and the final delivery. A temporary cursor holds only a pointer to that document, the approved revision, progress, strategy and evidence, and it is deleted when the work is done. The permanent document is the roadmap while the route is still unclear, and decisions depend on a frontier derived from it, which is an interesting inversion: the plan is the source and the frontier is the derived view, not the other way round. And the tool prefers an existing epic, request-for-comment or initiative document if the project already has one, creating a new directory only when it does not.
Editorial conclusion
CodeStable is worth reading if you are running AI-assisted development on something that will still exist in two years and are frustrated by the pile of plan files, status files and summary files that agent workflows leave behind, because its position is that the diff and the test output are the evidence and everything else is optional. It is a poor fit if you want coordination between several agents, since the project says plainly that it is not that and points you elsewhere for it, and a poor fit for a prototype, which the documentation also rules out by name. If you adopt it, cap your always-loaded context the way it does, insist that mechanically fixable mistakes become assertions rather than prose, and read the upgrade guide before moving from the first major, because the migration asks you to delete two dozen stale entry points by hand.
Frequently asked questions
What is CodeStable and what does it explicitly not do?
It is a set of skill contracts for human-in-the-loop AI coding. It states that it is not a multi-agent orchestration platform, not a process engine forcing every task through one pipeline, not a second system replacing existing documents or decision records, and not a necessary dependency for one-off prototypes.
What evidence does CodeStable require for different kinds of task?
A bug fix must go red then green, a refactor must first establish evidence of behaviour equivalence, and a feature needs design and validation proportionate to its risk. Independent review is performed by a read-only reviewer created from outside the task, in a single round.
How much project context does CodeStable load in every session?
At most twenty-five items, held in a single file created by onboarding. The cap is stated as a hard maximum, which is the design: anything read every session competes with the task, and an uncapped fact list degrades into noise.
How does CodeStable handle lessons, and what is the rule about mechanical errors?
Lessons live one per file and move through observed, validated and retired states, with duplicates checked and merged before writing. Creation is limited to at most one evidence-backed candidate per wrap-up and needs explicit authorisation. Errors that can be mechanised go into a test or a checker first rather than into a lesson.
What has to be done to upgrade a project from the first major version of CodeStable?
The upgrade guide asks you to precisely delete twenty-four retired entry points before installing the second major version, because installing on top would leave the old ones in place. A compatibility alias is kept for one retired review entry point so that existing invocations still resolve, but it only forwards and contains no rules of its own.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/codestable-codestable)