Zeroshot: an executor-verifier agent loop for code you can trust in prod
Your autonomous engineering team in a CLI. The agent loop produces senior-level code that you can actually trust in prod because of non-negotiable feedback from independent reviewers. Supports Claude Code, OpenAI Codex, OpenCode, and Gemini CLI with trivial setup.
At a glance
- What is it?
- Zeroshot is a CLI that drives coding agents through an executor-verifier loop, writing every step to a crash-safe SQLite ledger. It is for engineers who want autonomous changes but do not trust the agent that wrote the code to grade its own work.
- Who is it for?
- Adopt Zeroshot if you already have a provider CLI installed and want verified changes in an isolated worktree rather than a single agent's self-assessment. Do not adopt it if you are on Windows, since that platform is deferred, or if you only need trivial edits, because the TRIVIAL path runs a single worker with no validator.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 5 days ago.
- What is it written in?
- Mainly JavaScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem Zeroshot solves: self-graded agent output
The README opens with a line that doubles as the project's thesis: "The agent that wrote the code shouldn't be the one that says it works." That is the whole design constraint. Most coding-agent tooling has one model write a change and then either stop or ask the same model whether the change looks right. Zeroshot separates those two jobs into different agents with different context.
The audience is narrow and specific. This is for engineers who already run a provider CLI such as Claude Code, OpenAI Codex, OpenCode, or Gemini CLI, and who want an autonomous pass over a repository without hand-reviewing every line. It is not a chat interface and not an autocomplete layer. The unit of work is a task description, and the output is a verified change in an isolated workspace.
One detail in the README is worth flagging early because it undercuts the pitch slightly. TRIVIAL tasks route to a single-worker workflow with no validator at all. So the executor-verifier split, the thing the project is named around, does not apply to the easiest work. That is a reasonable cost decision, but it means the trust story is conditional on the conductor's classification being right.
How the executor-verifier loop actually runs
A conductor sizes the workflow before any code is written. It scores the task on complexity (TRIVIAL, SIMPLE, STANDARD, CRITICAL) and type (INQUIRY, TASK, DEBUG), and that score selects a workflow. A junior model runs the classification pass; if it cannot decide, it answers UNCERTAIN and a senior model takes over.
The rule table is evaluated top down and the first match wins. DEBUG above TRIVIAL always goes to debug-workflow, which runs an investigator, a fixer, a tester, and a completion-detector. TRIVIAL TASK or DEBUG with --pr or --ship goes to worker-validator. A bare TRIVIAL goes to single-worker. SIMPLE goes to worker-validator. STANDARD goes to full-workflow with a planner, a worker, and two validators. CRITICAL adds a meta-coordinator and four validators in two stages.
The part that matters for trust is the context boundary. According to the README, validators do not share the executor's session or reasoning context. They may receive explicit handoff artifacts, and they must reproduce reported failures. The loop runs until the change is verified or until it returns a concrete reason it is not.
Every step lands in a crash-safe SQLite ledger. That is the operational payoff: a run that dies mid-loop leaves a record you can inspect rather than a lost session. Provider keys are not stored by Zeroshot; it orchestrates the provider CLIs you already have configured.
Installing Zeroshot and running a first task
The README gives a two-command install. Node 22 or newer is required, plus one supported provider. Guided setup detects which providers are installed, picks a default, and configures worktree isolation for fresh repositories. Linux and macOS are supported today; Windows is deferred.
npm install -g @the-open-engine/zeroshot
zeroshotRunning zeroshot with no arguments starts guided setup. What you should see is a detected provider list and a prompt to choose a default. Once that finishes, move into a repository and hand it a task.
cd your-repo
zeroshot run "Add a --json flag with tests"In a git repository the guided default runs in a separate worktree, so your current checkout is not edited. The README is explicit that --no-isolation should be used only when you actually want the run to modify the current checkout. That default is the right one, and overriding it removes the main safety property.
To watch progress, open a second terminal. The list command shows runs and the logs command follows one by id.
zeroshot list
zeroshot logs <id> -fBefore running anything real, confirm which providers are visible and what the default is. These two commands are cheap and tell you whether setup did what you expect.
zeroshot providers
zeroshot providers set-default codexThe provider registry covers Claude, Codex, a bundled Gateway, Gemini, OpenCode, Pi, OMP, Kiro, and Copilot. Per-run override is available with --provider, for example zeroshot run 123 --provider gemini. A run can also be pointed at an issue instead of a string, and the README states that issue sources are auto-detected from repository context or explicit URLs across GitHub, GitLab, Jira, Azure DevOps, and Linear, each requiring its own authenticated client.
Custom workflows are JSON, and the bus underneath is the real interface
Each workflow in the routing table is a JSON file under cluster-templates/base-templates/, and the README states that none of them is privileged. Underneath is a message bus: agents subscribe to topics, publish to topics, and the graph is that wiring. Agent ids, roles, and topic names are free strings, and a trigger can carry a JavaScript predicate that decides whether a message wakes its agent.
That design has a real consequence. Cycles are legal, and reject-and-retry is one of them, but zeroshot config validate fails a ring of three or more unless something in the ring carries escape logic. Sub-clusters nest five deep. If you write a workflow that loops without an exit condition, the validator is the thing that catches you, not the runtime.
The inspection commands are the ones to learn first.
zeroshot config list
zeroshot config show full-workflow
zeroshot config validate ./mine.json
zeroshot run 123 --config ./mine.jsonMy read is that this is the most interesting part of the project and the least documented in the README. The routing table tells you what ships, but the message-bus model is what determines whether a custom workflow behaves. Anyone planning to modify a template should read an existing one end to end before writing a new one, because the trigger predicates are where the control flow actually lives.
Isolation, delivery flags, and where the loop breaks down
Isolation modes cascade through delivery flags. The README states that --ship implies --pr, which implies --worktree. Git worktree isolation gives you an isolated branch and checkout and is the guided default for fresh repositories. Docker isolation is available via --docker for what the README calls riskier workloads.
The clearest limitation is platform support. Windows is deferred, stated plainly in the install section, with no timeline given. If your team develops on Windows, this is not a tool you can adopt today regardless of how the loop is designed.
A second limitation is classification risk. The conductor is instructed to pick STANDARD whenever it is torn, because CRITICAL spends a senior model and four validators. That instruction protects against over-spending, but it means an under-classified task gets a lighter workflow than it deserved. A task misread as TRIVIAL runs one worker and no verifier, which is exactly the failure mode the project exists to prevent.
A third is the dependency surface. Zeroshot orchestrates provider CLIs rather than shipping its own model access, so a broken or unauthenticated provider CLI breaks the run. The README says each issue source requires its own authenticated client where applicable. There is nothing in the README about rollback behaviour for a partially applied run, so treat that as unverified and check the ledger before assuming a clean state.
There is also a second product in the same repository. Zeroshot Rust is described as an independent native product with its own zeroshot-rust CLI, its own releases, and a self-hosted target image. The README states that the remainder of the document describes the established Node product, so do not assume the two share flags or configuration.
Zeroshot compared with a single-agent CLI
The obvious alternative is running Claude Code, Codex, or Gemini CLI directly and reviewing the diff yourself. The difference is not capability, since Zeroshot uses those same CLIs as providers. The difference is who judges the result.
With a single agent, verification is either your review or the same model's self-assessment in the same session. With Zeroshot, a separate verifier with no shared reasoning context must reproduce reported failures before the loop closes. That is a structural separation, not a prompt instruction, and it is the reason the project can claim the executor is not the one saying it works.
The trade-off is cost and latency. STANDARD spends a planner, a worker, and two validators. CRITICAL spends a senior model and four validators across two stages. A single-agent CLI spends one model pass. If your tasks are small and you review diffs quickly anyway, the loop is overhead you will notice on every run.
The other alternative is a general CI pipeline with test gates. That catches regressions but cannot investigate a failing behavior or decide whether a change satisfies an open-ended task description. Zeroshot's debug-workflow, with an investigator, fixer, tester, and completion-detector, is aimed at that gap. If your work is already well covered by deterministic tests, a CI gate is cheaper and more predictable.
Maintenance, licensing, and what a run costs you to keep
The repository is not archived, and the last push was on 2026-08-28. Recent releases include zeroshot-rust-v0.5.0 on that same date, plus v6.45.0 and v6.44.0 on 2026-08-26. The project is MIT licensed at the repository level, and the Cargo workspace declares license = "MIT" as well. MIT is permissive, so the practical implication is that you can use and modify it in commercial settings, but the software ships without warranty. That is a statement about the licence text, not legal advice; check with your own counsel if the distinction matters to your organisation.
Upgrade cost is the part worth thinking about before adopting. The repository carries two products with separate release streams, a Node CLI and a Rust CLI, and the README treats them as distinct. The Node CLI is the established one. If you build workflows as JSON files under cluster-templates/base-templates/, those files are yours to maintain, and the validator will fail them if the graph contains an unescaped cycle. Tooling in package.json runs protocol:check and distribution:check alongside lint and tests, which suggests the protocol surface is versioned and checked rather than incidental.
Provider churn is the ongoing cost that is easy to underestimate. Zeroshot orchestrates external CLIs, so a breaking change in a provider's interface lands on your runs. The registry lists nine providers, and each is a moving dependency you do not control.
Editorial conclusion
Adopt Zeroshot if you already have a provider CLI installed and want verified changes in an isolated worktree rather than a single agent's self-assessment. Do not adopt it if you are on Windows, since that platform is deferred, or if you only need trivial edits, because the TRIVIAL path runs a single worker with no validator. Before committing, verify that the provider you intend to use is detected by the guided setup, check whether worktree isolation is configured for your repository, and read cluster-templates/base-templates/full-workflow.json to see what the loop will actually do on a STANDARD task.
Frequently asked questions
What is Zeroshot?
Zeroshot is a CLI that drives coding agents through an executor-verifier loop, where a separate verifier judges the observable result without sharing the executor's session or reasoning context. It supports Claude Code, OpenAI Codex, OpenCode, and Gemini CLI, among other providers.
How do I install Zeroshot?
Install it globally with npm and then run the bare command to start guided setup. It requires Node 22 or newer and one supported provider, and guided setup detects installed providers and configures worktree isolation for fresh repositories.
Does Zeroshot modify my current checkout?
By default it does not. In a git repository the guided default runs in a separate worktree, and the README says to use --no-isolation only when you explicitly want the run to modify the current checkout.
Which platforms does Zeroshot support?
The README states Linux and macOS today, with Windows deferred. No timeline for Windows support is given.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/the-open-engine-zeroshot)