Devora's bet is that the missing context, not the missing code, is what an agent gets wrong
AI agent for end-to-end software development, from requirements to operations.
At a glance
- What is it?
- A control plane that sits on top of a coding agent rather than replacing it, built around three nouns: what is true about the project, how recurring work is done, and what is allowed within one piece of work. Two runtime dependencies.
- Who is it for?
- Two things to check before you adopt it. The first is whether your agent already reads your project documentation, because if it does, this adds a structure rather than filling a gap, and that is still worth having but is not the same pitch.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 34 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.
Editorial analysis
It sits on top of your agent rather than replacing it
The framing is a control plane and project context for coding agents, and the second paragraph explains why that exists. Agents can read code. What they cannot reliably do is infer every business boundary, engineering rule, system relationship and historical decision.
The stated consequences of that gap are three: you explain the same thing repeatedly, implementations drift apart from each other, and changes become unsafe. All three are process failures rather than capability failures, which is why the answer here is not a better model.
The positioning is explicit and it is the most honest paragraph in the readme. It sits on top of Codex, Claude Code, an open-source editor agent, an editor, or another coding agent. It does not replace the agent, does not run a model of its own, and does not automatically rewrite problems it finds.
That last clause matters more than it sounds. A tool in this category that silently refactors your codebase is a tool nobody can put in a review process. One that only produces structure and lets the agent act is something you can stage.
Install is two commands and a runtime floor of a recent Python, with a package installer doing the work rather than pip:
uv tool install devora-cli
cd /path/to/project
devora init . --integration codex --language enThe last two flags are the interesting part of that command. You name your agent and your working language at init time, and the tool writes a matching entry point for it. Four of the supported agents use a dollar-prefixed invocation and the rest use a slash command, and the readme tells you which is which so you do not spend an hour wondering why nothing happened.
Three nouns, and the whole system is the relationships between them
There are exactly three concepts, which is the sign of a designer who has decided what matters.
Governance is what is true and what is allowed: architecture, engineering rules, business constraints, risks, decisions. A Skill is how recurring work should be done: steps, constraints, scenarios and validation evidence. A Change is one piece of work from requirement to review: scope, source material, tasks, tests, approvals and handoff.
The relationship between them is the design. Governance defines the constraints, Skills encode recurring procedures inside them, and each Change applies both to a specific task and then returns what it learned. That last clause is the one that makes the system compound rather than accumulate, and it is stated explicitly: confirmed decisions, rules, risks and reusable workflows go back into the project context instead of disappearing into chat history.
Compare that with what usually happens. An agent session produces a good solution, you review it, you merge it, and six weeks later another agent has no idea that the first one discovered the payment provider has a rate limit. The learning is the expensive part and it is almost always thrown away.
The six-step lifecycle follows directly: initialise, understand and clarify, reuse or build skills, develop inside a controlled change, validate and review, and feed durable conclusions back. That last step being part of the loop rather than an afterthought is what makes it a system rather than a checklist.
What lands on disk, and what deliberately does not
The generated structure is small enough to read in one screen, which for a governance tool is the point:
.devora/
├── project.md # Entry point and current project understanding
├── context/ # Durable project knowledge
│ ├── product.md
│ ├── engineering.md
│ ├── architecture.md
│ ├── rules.md
│ ├── risks.md
│ └── decisions.md
├── skills/ # Project-specific Skills
├── changes/ # Current work and compact history
│ └── history/
└── state.json # Machine-readable workflow stateSix context files, a skills folder, a changes folder with a history subfolder, and one machine-readable state file. Human-readable knowledge is markdown; workflow state is JSON. That separation is deliberate and stated: you can read the context without the tool, and the tool can be replaced without losing the knowledge.
Then comes the part most tools of this kind get wrong, which is the list of things it refuses to create. No separate roles directory. No custom directory. No project-local scripts directory. No integrations directory. Four omissions, each of which would be an obvious place for a tool to accumulate project-specific machinery until nobody can predict what it does.
The Change structure is graded by complexity rather than fixed. A routine one is a single file: a directory named for the change containing one document. A complex or cross-repository one adds a task list, a review document, an evidence folder and a handoff folder, on demand. Completed changes are compacted into a history folder.
That compaction rule is worth naming because it is the difference between a tool that helps after six months and one you abandon. A change system without compaction becomes a tree of nested directories that grows forever, and once that happens the tool is overhead rather than leverage.
Two risk taxonomies, and why quality gates are not the test suite
The safety model has two axes and they are not the same thing. Technical risk covers database mutation, permissions, infrastructure, compatibility, concurrency, production operations and recovery. Business risk covers payments and assets, privacy, pricing, inventory, user rights, compliance, approval and bulk operations.
Two axes because they need different reviewers. A senior engineer can tell you whether a migration is safe. Only someone in the business can tell you whether a refund flow touches something contractual.
Then the rule that matters, and it is the one sentence I would put at the top of an evaluation: a passed unit suite does not override a missing security check, an unresolved business rule, a stale contract, or a required approval.
That is a direct attack on the most common failure in agent-assisted development. An agent runs the tests, the tests pass, and the merge happens. The tests were always going to pass, because the agent wrote them. The unit suite is the weakest possible evidence when the code under test was generated in the same session.
The other capabilities in this section are what a real review process needs and an agent does not have: freeze the authorised paths, so the implementation cannot escape its approved scope; track requirement sources and cross-repository contracts; discover what test capabilities the project already has; record repeatable evidence rather than a claim that something was checked; and audit afterwards whether the change stayed inside its boundary.
The last one is the unusual feature. Most tooling checks before. This one also checks that what came back did not do more than it was asked.
Requirement sources are first class, and conflicts block work
One of the example invocations is worth reading closely, because it demonstrates the model better than the feature list does. Implement a refund flow, using three different sources of truth: a product requirements document at a URL, a design image on disk, and an interface definition file.
All three of those are things a real project has and none of them is a code file. The image in particular is the interesting one: a mockup in a design tool is often the most authoritative specification a team has, and it is not something an agent can read without a tool that understands images.
The rule attached to it is the one that makes this more than a note-taking tool. Material conflicts and source drift block development until they are resolved or explicitly accepted. Not warn: block.
That is a strong stance and a defensible one, because the failure it prevents is the expensive one. An agent given a requirements document and an interface definition that disagree will pick one, and it will pick silently, and you will find out during integration.
The other two modes are about not doing too much. A validation-only change runs the release checks without touching product code, and the reason given is that findings should still receive explicit review without silently expanding into remediation. A governance-only change updates knowledge without implementing anything.
Both exist because the default failure of an agent asked to check something is that it fixes it. Constraining the change type is a way of expressing that some tasks have a scope and you do not want the scope renegotiated.
Two runtime dependencies, and what that implies about where it will break
The package manifest is short enough to be a design statement. Two runtime dependencies: a command line framework and a terminal formatting library. Everything else is test tooling, and the tests themselves are two packages.
That is remarkable for a project of this scope. No framework, no template engine, no database, no agent SDK, no model client. The readme is explicit that it does not run a model of its own, and the dependency list confirms it.
The consequence is worth understanding, because it cuts both ways. On the good side: nothing here can break because a vendor changed an API, and the tool will still work in five years as long as those two libraries exist. You can read the entire dependency surface of the tool in one screen.
On the other side: everything the tool does is text processing and file management, which means the intelligence is in the agent, not here. If you point it at an agent that ignores its instructions, this tool produces a directory structure and nothing else.
So the value is entirely in the structure being followed. That is a real risk and it is worth being clear-eyed about: this is a protocol that the agent is asked to follow, not a mechanism that enforces it. Nobody is verifying that your agent actually froze the authorised paths.
The packaging itself is conventional and well done. A declarative build backend, a wheel target that includes the package, and a second forced include that ships the templates directory inside the installed package. That last line matters: template files that are not packaged are the most common packaging bug in this kind of tool, and it is handled explicitly.
The test configuration is minimal and disciplined: strict markers on, quiet output, one test path. Strict markers means an unregistered marker is an error rather than a warning, which is the setting you want in a repository where slow and integration tests are marked.
Editorial conclusion
Two things to check before you adopt it. The first is whether your agent already reads your project documentation, because if it does, this adds a structure rather than filling a gap, and that is still worth having but is not the same pitch. The second is the directory discipline. This tool's value depends entirely on the knowledge files being accurate and being maintained, and a stale rules file is worse than none, because the agent will follow it confidently. That is also why the design puts the rules in your repository under version control with review: the failure mode of this approach is not that it does not work, it is that it works perfectly on out-of-date knowledge.
Frequently asked questions
What does Devora do?
It is a control plane that sits on top of a coding agent, keeping project context in versioned files and putting each piece of work inside an explicit boundary. It does not replace the agent, does not run a model of its own, and does not automatically rewrite problems it finds. You install a command line tool and then invoke it from inside your agent.
What are the three core concepts?
Governance, which is what is true about the project and what is allowed; Skills, which encode how recurring work should be done; and a Change, which is one piece of work from requirement through review. Governance sets the constraints, Skills encode procedures inside them, and each Change applies both and then feeds what it learned back into the project context.
What files does it create in my repository?
A small workspace with one entry-point document, six markdown context files covering product, engineering, architecture, rules, risks and decisions, a skills folder, a changes folder with history, and one machine-readable state file. Human-readable knowledge stays in markdown and workflow state stays in JSON, so you can read the knowledge without the tool.
How is this different from just writing good documentation?
The differences are structural rather than editorial. Requirement sources are tracked, and material conflicts between them block development until resolved rather than warning. Each piece of work has a frozen set of authorised paths and is audited afterwards for escaping its scope. Quality gates are selected by risk, and a passing unit suite explicitly does not override a missing security check or a required approval.
Which coding agents does it support?
Several, with the integration named at initialisation time along with your working language, so the tool writes a matching entry point. Some use a dollar-prefixed invocation and others a slash command, and the readme specifies which is which per agent. The tool sits on top of whichever agent you already use rather than bundling one.
What does it require to install?
A recent Python and a package installer rather than pip; the install command is a tool install followed by an initialisation command in your project. The runtime dependency list is two packages, a command line framework and a terminal formatting library, which means the tool itself is text and file processing and depends on your agent for its intelligence.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/cheney369-devora)