claude-dynamic-workflows-codex, one slash command that can launch a fleet
Run Claude Code dynamic workflows on a local Codex (GPT) backend, plus an interactive run viewer
At a glance
- What is it?
- This project takes a workflow language that ships with one coding agent and runs it against a locally installed backend from another vendor, with every worker pinned to the strongest available model. The entry point is a single manual slash command, and the readme is explicit that it never fires on its own. What that command can do is the interesting part: a flag turns one workflow into a fleet of concurrent workflows that the first agent supervises, answering their gates and cancelling the ones that lose, with no cost ceiling mentioned anywhere.
- Who is it for?
- Use this if you want to fan a task out across many workers on a backend you run yourself, and you are the kind of user who wants a live execution map rather than a single answer. Four things to know before the first run.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 88 days ago.
- What is it written in?
- Mainly JavaScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 7, 2026, and from our analysis. They are not legal advice.
Editorial analysis
A workflow language built for one model runs on another
The primitives are five: one that starts an agent, one that runs them in parallel, one that chains them, one that groups work into phases, and a budget. That is the whole language, and it is the language that ships with the other agent.
Running it on a different vendor's backend is the project's reason to exist, and the readme does not hedge about it. What it does say is how the model is chosen, and the answer is that every worker in a run is pinned to the current flagship of the other backend, detected dynamically at run time. The readme is explicit that it does not route stages across the cheaper tiers of that model's family, even though three tiers are named and available.
So a pipeline with a cheap gather stage and an expensive judge stage runs both at the flagship price. The readme's own captured output shows the shape of the bill: a six-agent, two-phase run consuming just over seven hundred thousand tokens over twenty minutes.
The language is portable in syntax and the model policy is not, and the readme does not discuss which of the five primitives depend on the authoring agent's behaviour rather than the language itself.
One flag turns a workflow into a fleet with no stated ceiling
There is a base mode and there is a flag. The base mode is already wide: the readme describes the workflow running across dozens of workers, with the runtime holding the loop and the branching so that only the final answer reaches the conversation.
The flag changes the level rather than the width. With it, the first agent launches a whole fleet of concurrent workflows and supervises them itself. The readme lists what that supervision consists of: answering the gates those workflows are blocked on, steering them, killing dead ends, and forking the winners.
That is the mechanism worth reading twice, because the gates are the human-decision points of the inner workflows and the outer agent answers them. So a command a person types once can produce a tree of agents, each of which can fan out again, with decisions made at the leaves by something that is not the person.
The language has a budget primitive, so the mechanism for bounding cost exists. The description of the fleet flag does not mention it, does not give a default, and does not say what happens when a budget is exhausted. The captured output gives the only cost figure in the readme, and it is for a single six-agent run rather than for a fleet.
The skill never fires on its own, and that is the only bound on the fan-out
The readme states one safety property plainly: the skill is for manual invocation only, and the agent never triggers it by itself.
That sentence is doing a lot of work, because everything else in the document is about scale. The base mode fans out across dozens of workers. The flag multiplies that into concurrent workflows. Those workflows have gates that block waiting for an answer, and in fleet mode the answer comes from the supervising agent rather than from a person.
So the entry point is bounded and the consequences are not. Nothing in the readme describes a cap on the number of workflows, a cap on the number of workers per workflow, a default budget, or a dry run. The word manual appears in the sense of a human typing a command, not in the sense of a human approving each step.
For a user who understands the flag, that is a design they have opted into. For a user who reads the capability list without the safety sentence, the two facts are in different sections of a long document, and the safety sentence is the shorter one.
The recommended install tracks the branch and the only release is older than the code
There are two install paths and the readme recommends the first for a specific reason: it updates with every push.
That reason is also the risk. The recommended path adds a plugin marketplace entry pointing at the repository and installs the skill from it, so the installed skill is whatever the default branch contains at the moment you install or update. There is no version to pin and no way to say which commit you have.
The alternative is a clone into a dot directory in your home folder, which gives you a fixed commit and a normal git workflow. The readme also documents a script that pushes the skill surface from a working clone into that same home directory, so there are three ways the skill files arrive there and the third overwrites whatever the first two put there.
The release history explains why none of this matters much yet. There is exactly one release, tagged in June 2026, and the branch was last pushed in July 2026, after it. The manifest also marks the package as not publishable, which is why the readme invokes it through a source specifier rather than a registry name.
So there is one tag in the project's life, the code has moved past it, and the recommended install ignores tags entirely. For a project at version 0.2 that is a defensible choice, and for anyone who needs reproducibility it is the wrong default.
The documented way to run it executes whatever the default branch holds
The documented way to verify an install is one command, run through a source specifier that names the repository:
npx github:scasella/claude-dynamic-workflows-codex doctor # → state: readySo the install is two plugin commands or a clone into a dot directory in your home folder, and the verification is a package runner pointed at a git reference. There is no registry copy, because the manifest marks the package as not publishable, which is why the invocation looks like that rather than like a package name.
The phrase without installing anything is accurate about the install step and misleading about the code. A source specifier resolves the default branch, fetches it, and runs it. So the health check offered as the way to verify your setup is running code from a moving reference, and so is every subcommand after it.
The check itself is worth noting. It is described as a handshake, its success output is a single word followed by a state, and nothing is given as an example of what it prints when the backend is missing, not logged in, or on the wrong path.
The health check is a file that lives in the test directory
The manifest has a script called doctor, and its entire body is a path into the test folder. The same is true of several other scripts: the demo, the viewer, the mapper, the fleet tool, and the workflow runner are all files under the runner's binary directory invoked directly rather than through declared executables.
So the diagnostic a new user is told to run is a test file, and the test suite and the command line tools share a directory. That is a common shape for a small project and it has one real consequence: there is no boundary between a check that reports a problem and a test that fails.
The output contract is thin as well. The readme quotes a successful result as a state and one word. There is no documented output for the failure cases a new user will actually hit: the backend command not on the search path, the login not done, the model name not detected, or the version of the runtime below what the tools need.
The prerequisites section lists two things, a Node version floor and a logged-in backend command, and offers one command to verify both. For a tool whose first documented step is compiling a workflow script and launching dozens of workers, one word of confirmation is thin.
Fifteen test files chained with && and one of them is live
The test script is one long chain. Fifteen file paths joined by a logical and, in a fixed order, with no test runner involved.
Three things follow from that shape. A failure at position three means positions four through fifteen never run, so a run tells you about one problem rather than about the state of the suite. There is no filtering, so the only way to run a subset is to name a file by hand on the command line. And there is no coverage configuration anywhere in the manifest, so nobody is measuring what the suite reaches.
The chain also mixes categories. The first entry is a file that does not carry the test suffix the other fourteen do, so it is either a shared offline harness or a test that was renamed. The fifth entry carries a live suffix, which means the default test command expects a working backend connection and a real model, not a mocked one.
For a project whose entire value is orchestrating a remote model across a protocol, having a live test in the default chain is defensible, and having no offline-only variant of the command is not. The readme does not mention a test command at all.
Twenty curated workflows in a directory and two in the repository root
The examples directory holds twenty named workflow files and four subdirectories. The names are a good index of what the language is for: a bug hunt, a code review, a review with gates, a deep research run, a classification pass over a route, a flaky test perturbation experiment, a tournament sort, a stateful dialogue, a sessionful workers example, a warm context interrogation, a lead following research run, an agent foreman, and a hedged take-first-win pattern.
The subdirectories are benchmarks, the bundled demo, fleet examples, and a collection of harnesses. So the project ships a zoo of reusable task shapes alongside the specific workflows.
Then there are two workflow files in the repository root, next to the readme and the licence, rather than in the examples directory with everything else. Their names say they are for polishing the text output of the map view and for polishing the graphical view.
That is the project using its own tool on its own interface, which is the best kind of dogfooding available in a repository like this. It is also two working files in the root, which is where a readme, a licence, a contributing guide and a plugin manifest belong.
Editorial conclusion
Use this if you want to fan a task out across many workers on a backend you run yourself, and you are the kind of user who wants a live execution map rather than a single answer. Four things to know before the first run. That every worker is pinned to the strongest model, so cost scales with fan-out width and there is no tier routing to keep a cheap stage cheap. That the language was authored for a different agent's model, and the readme does not say which of its primitives behave differently on the other backend. That the recommended install tracks the branch rather than a release, so what you run today is whatever the default branch contains. And that the health check the readme points you at is a test file whose success output is one word and whose failure output is not documented.
Frequently asked questions
What are Claude Dynamic workflows?
They are a small orchestration language with five primitives: one that starts an agent, one that runs several in parallel, one that chains them, one that groups work into phases, and a budget. A runtime holds the loop, the branching and the intermediate results, so the conversation only sees the final answer, and the run is shown as an interactive map rather than a single reply.
What is Codex and how is it different from Claude in this project?
In this project Codex is the local backend the workflows run on, reached through its own command line tool, while Claude Code is the agent that authors the workflow script and supervises the run. The readme states the project is unofficial and not affiliated with either vendor, and that both names are trademarks of their respective owners. Every worker is pinned to the current flagship model of the Codex backend rather than spread across its cheaper tiers.
How much does a claude-dynamic-workflows-codex run cost?
The readme does not give a figure. It shows one captured run of six agents across two phases consuming just over seven hundred thousand tokens over about twenty minutes, with every worker on the flagship model and no routing to cheaper tiers. The language has a budget primitive, but the description of the fleet flag does not mention a default budget or a cap on the number of concurrent workflows.
How do I install and verify claude-dynamic-workflows-codex?
Either add the repository as a plugin marketplace to the coding agent and install the skill, which the readme recommends because it updates with every push, or clone the repository into a skills directory in your home folder for a fixed commit. You need Node 18 or newer with no packages to install, and the backend command line tool on your search path, logged in. A single command then reports a ready state.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/scasella-claude-dynamic-workflows-codex)