codex_workflow's benchmark section promises a token saving and contains none
A swarm orchestration system in Codex - use less quota, get more done.
At a glance
- What is it?
- viettran-edgeAI/codex_workflow is a swarm orchestration system for the Codex CLI that delegates work to subagents across three routes. Its central claim is token efficiency, and its own advice is to use the highest reasoning effort rather than a low one. The section headed as a benchmark introduces a cost-saving technique, ends in a colon, and shows no figures at all.
- Who is it for?
- Use it if you run long agentic tasks where the main agent's own turn count is what costs you, since that is the mechanism it targets and the batching rules are specific. Do not adopt it expecting a measured saving, because the one section that would carry that number is empty and the repository calls its own measurement an initial case study with a link to a proposal for testing it properly.
- Can I use it commercially?
- Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
- Is it still maintained?
- Yes. The repository last received commits 6 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The benchmark section ends in a colon and shows no figures
The most important section of this repository is the one with nothing in it. It is headed as a light benchmark, and its only sentence says that a batching technique introduced in version 1.1.3 significantly reduces the main agent's rollouts, which in turn reduces the main agent's cached input tokens, described as a major component of operating cost. The sentence ends with a colon.
After the colon there is no table, no chart, no number and no comparison. The next paragraph says the current benchmark is an initial case study and points the reader at a coverage proposal document for ideas on testing more tasks and more model providers.
That proposal link is the honest part. It says out loud that one case study is not a benchmark and describes what would be needed to test it further. But the practical effect is that the product's headline claim, use less quota, has no published measurement in the repository, and the section that would carry it reads as a stub that was written and then abandoned.
The mechanism the project describes is plausible and specific: the main agent's turn count is what drives cached input cost, batching collapses several worker results into one main-agent turn, and fewer main-agent turns means less repeated context. What is missing is any indication of how much that saves, on what task, against what baseline.
The token-saving tool tells you to use the highest reasoning effort
The recommendation is the opposite of what the product name implies. It says to assign very large and complex tasks to the heavy route and not to hesitate when choosing the highest available reasoning efforts, on the grounds that this maximises token savings. It then adds the sentence that makes the claim coherent: using much lower reasoning efforts will not actually save tokens, and will severely reduce the system's coordination capabilities.
So the position is that the saving comes from architecture rather than from thinking less. The main agent's own turns are batched, so fewer cached-context turns are billed, and within each turn the workers can afford to reason more deeply. Under that model, dropping the reasoning effort does not reduce cost; it just produces worse plans that need more main-agent turns to recover from.
That is a defensible and fairly interesting claim. It is also, notably, an argument rather than a measurement, and it appears in the same repository as the empty benchmark section.
A second mechanism is offered for the same goal. Worker reports are supposed to preserve the material evidence while referencing bulky logs and artefacts instead of copying them into the result, and a batching guideline is described as preventing excessive main-agent rollout. The design consistently tries to move bytes out of the main agent's context rather than to make any individual step cheaper.
The medium route is documented as costing more than the heavy route
Three routes are offered, and their cost relationship does not follow their size.
The light route is the default and turns off workflow mode entirely with minimal context. The medium route adds three read-only subagents, one for discovery, one for solution research and one for documentation, while the main agent keeps implementation and verification. The heavy route delegates bounded production and verification to executor and tester subagents, keeps the same three support subagents, and leaves the main agent with orchestration, synthesis and decisions only.
Medium therefore does less delegation than heavy, which you would expect to be cheaper. The documentation says the opposite in as many words: choose medium when you want context support without delegating production work, such as front-end design, visualisation or three-dimensional work, and it will burn tokens faster than the heavy route.
The reason is implicit rather than stated. In the medium route the main agent still does the production work itself, so you have paid for three subagents' context and kept the main agent's expensive turn count as well. In the heavy route the main agent stops doing production work, so its context stays small enough that batching pays off.
That makes the ordering coherent with the batching thesis, and it is worth knowing before choosing the middle option, since the intuitive pick is the expensive one.
Installation is a prompt you paste, and it asks for full access first
There is no installer script to run and no package to add. Installation is a natural-language instruction that you send to the coding agent from inside your project directory:
Download and extract the latest `codex_workflow-<version>.zip` asset from https://github.com/viettran-edgeAI/codex_workflow/releases. Verify it against `SHA256SUMS`, then read the bundled `codex_workflow/operate/bootstrap.md` and follow it to complete the initial installation.The agent downloads the archive, checks it against a checksum file, opens an instruction document bundled inside the archive it just downloaded, and follows that. The documentation asks you to change the agent's permission setting to an approve-for-me or full-access level before sending the prompt, and to restart the agent afterwards. The step word is misspelled in the original, which is a small signal that this section has not been read carefully.
Once installed, subsequent projects use a command, and the first bootstrap is said to create the project's documentation directory using a documentation subagent. Python 3.11 or newer is required, with the reason given as deterministic lifecycle operations.
The trust model is worth stating plainly: you are asking an agent with elevated permissions to fetch an archive from the internet and obey instructions contained inside it. That is a category of risk that a package manager with a lockfile does not present.
The checksum protects against corruption, not against a bad release
Requiring verification against a checksum file is good practice and the instruction is explicit about where it happens. What the instruction does not say is where the checksum file comes from.
As written, the agent is told to download an asset from the release page and verify it against a file that is, by implication, obtained from that same release page over the same connection. That protects against a truncated download and against a mirror serving something different. It does not protect against the case that actually matters here, which is a release that has been replaced or an account that has been compromised: an attacker who can publish a bad archive can publish a matching checksum.
Closing that gap needs the checksum to arrive through an independent channel, a signed tag, or a published digest whose fingerprint a reader already trusts. None of those appear in the visible instructions.
The mitigating factor is that the trust model is already agent-mediated rather than automated, so a human is nominally in the loop. In practice the agent performs the fetch and the comparison, and the user is asked to trust both. The repository does include a releasing guide at the top level, which is where a signature or an out-of-band digest would most naturally be documented.
The workflow writes into your user-level agent instruction file
One of the design decisions is stated almost in passing and deserves more attention than it gets. Workflow policy is said to live in a managed user-level agent instruction region under the home directory, while each project's own instruction file remains native and project-owned personalisation.
That means installing this changes a file that applies to every project you open, not just the one you installed it into. The distinction the project draws is sensible in principle: global defaults for how the agent behaves everywhere, project-local files for your own preferences. In practice it means a tool installed once silently rewrites your global agent instructions, and the per-project story only covers the documentation directory.
The other durable side effect is that project directory. It is described as built-in project memory available on the two larger routes, holding project goals, architecture, progress, decisions and the latest handoff across sessions, with the framework instructions for the routes themselves living in those documents. That is the most valuable part of the system and the part most likely to be genuinely useful, since it is what survives a session boundary.
Alongside it the system reports its own token use at the end of each session, per agent, so you can see which role is consuming the budget.
Install is scoped to a project and update is scoped to you
Five commands are documented in a table, and the difference in scope between them is not obvious from the names.
The install command installs the workflow in the current project and initialises its documentation framework. The update command installs a newer release for the user and the current project, and it also does something else: it brings the current project up to a release that is already installed. So one verb covers both a machine-level upgrade and a project-level catch-up, and which one you get depends on state the documentation does not describe.
The remove command removes the installed workflow after a destructive dry-run and a confirmation, which is the right shape for a command that deletes files across a project. The check-update command reports whether a newer release exists without installing it, and the version command reports what is installed.
Two repository layout details round this out. There is a build output directory committed at the top level, which is not where a Python project usually keeps anything it did not compile, and three images sit at the root rather than in an assets directory, two of them diagrams of the routing structure and one a token report capture. Neither is wrong, and both suggest the repository is organised around documents rather than around a package.
Editorial conclusion
Use it if you run long agentic tasks where the main agent's own turn count is what costs you, since that is the mechanism it targets and the batching rules are specific. Do not adopt it expecting a measured saving, because the one section that would carry that number is empty and the repository calls its own measurement an initial case study with a link to a proposal for testing it properly. Before you install, understand what installation actually is: you paste a prompt into the agent, you raise its permission level first, and the agent then downloads a release archive and follows instructions from inside it. Confirm you accept that trust model, check where the checksum file comes from, and note that the tool writes into your user-level agent instruction file as well as your project.
Frequently asked questions
what is codex workflow
A swarm orchestration system for the Codex command line interface. It adds three routes over the agent's normal behaviour: a default lightweight route with no workflow mode, a medium route that adds read-only discovery, research and documentation subagents, and a heavy route that delegates production and verification to executor and tester subagents while the main agent keeps orchestration and decisions.
how to use codex workflow
For simple work and general questions you do nothing, since the light route is the default. For larger work, send the agent a prompt starting with use medium/heavy route followed by your task description, or followed by Continue ongoing work for a task from a previous session. The agent stays on the selected route until you change it.
How do I install codex_workflow?
From your project directory, change the agent's permission setting to an approve-for-me or full-access level, then send it a prompt telling it to download the latest release archive, verify it against the checksum file, and follow the bundled bootstrap instructions. Python 3.11 or newer is required, and the agent is restarted afterwards.
Does codex_workflow prove that it saves tokens?
Not in the repository. The section headed as a benchmark introduces a batching technique said to reduce the main agent's rollouts and its cached input tokens, and then contains no figures, with the text calling the current benchmark an initial case study and linking a proposal for testing more tasks and more providers.
Which route uses the fewest tokens in codex_workflow?
The documentation says the medium route burns tokens faster than the heavy route, even though the medium route delegates less. The stated reason is that on medium the main agent still does the production work itself, so you pay for the support subagents and keep the main agent's expensive turn count, while heavy stops the main agent doing production work so batching pays off.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/viettran-edgeai-codex-workflow)