foreman supervises the claude CLI you already have
A Boris-style agentic orchestrator TUI that supervises headless Claude Code agents through a gated software-delivery pipeline — pointed at any repository.
At a glance
- What is it?
- A Python Textual TUI that sits between you and headless claude CLI agents, running a gated pipeline from plan through e2e with human review on the design phases. Two runtime dependencies, no database, and guardrails the orchestrator enforces rather than trusting the agent to self-police.
- Who is it for?
- It fits someone already running claude in headless mode who lost an afternoon to an agent that edited past review and then approved its own work, because the hash-sealed approval and the deny hook target exactly that failure. It does not fit macOS or bare Linux users, since the stated support is Linux or WSL2 tested on Ubuntu, and it does not fit anyone who wants an agent library rather than a supervised workflow.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 103 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 10, 2026, and from our analysis. They are not legal advice.
Editorial analysis
It drives the claude binary you already have
It does not ship credentials. It spawns the locally installed claude CLI in headless stream-json mode, parses the event stream that comes back, and enforces budgets around each run. The requirement list is explicit that the claude CLI must be installed and authenticated, with claude --version named as the check, alongside git and Python 3.11 or newer. That design has a practical consequence: the agent's behaviour is whatever your local CLI does, and Foreman's leverage is entirely in supervision. Budgets are per run rather than global, covering turns, cost and time, so a runaway loop is bounded by the orchestrator instead of by the agent noticing. The framing in the documentation is that running a coding agent in a while loop is easy while running one you can trust to merge is not, and this spawn-parse-enforce loop is the whole mechanism behind that claim. Installation is a single command, either pipx install . or uv tool install ., which exposes one foreman command and nothing else. The documented path from there is short:
# 1. Install (exposes a single `foreman` command)
pipx install . # or: uv tool install .
# 2. Point it at any repo
cd /path/to/your/repo
foreman init # scaffolds .foreman/ and installs the foreman-* skills
# into .claude/skills/
# 3. See the whole thing work end-to-end with NO tokens spent
foreman demo # runs the full pipeline against a throwaway sample repo
# using a mocked agent backend (canned stream-json)The same page states the aim in a line worth repeating, that the author no longer prompts Claude directly but runs loops that prompt Claude, and the tool is described as pointed at any repository rather than tied to one project.
Approval is hash sealed and undoes itself when a doc changes
The pipeline is written as a chain: plan, then ADR and PRD, then issues, then a TDD build, then end-to-end. Human review gates sit on every design phase, and the mechanism underneath them is worth spelling out. Approval is hash sealed, meaning the gate records a hash of the document you approved, and if that document changes later the approval is treated as void and the phase reverts automatically. Without that property a gate is only a speed bump, since an approved plan can be edited after the fact and still carry its approval. The phases you see on the dashboard map to single keys: the request phase runs the planner, plan review opens with a to approve and r to request changes, grilling runs the ADR and PRD pass, doc review opens again, and slicing runs the slicer that turns documents into a queue. The TUI tells you the next key to press in its hint line, so the flow is driven by keyboard rather than by menus. Every screen shows its own keys in the footer, the secondary screens for review, workers, attention, metrics, retro and settings are pushed on top with a single key and dismissed with Esc, and the dashboard itself lists your features on the left with the selected feature's phase and cost on the right above a live issue board.
State is files in your repo, so a kill is recoverable
There is no database anywhere in the design. All state is written as human-readable files committed inside the target repository, which is what makes the crash-safety claim concrete: kill the process mid build and it recovers from disk, because the disk copy is the authoritative one. Pointing Foreman at a repository starts with foreman init, which scaffolds a .foreman directory and installs the foreman prefixed skills into .claude/skills, meaning the target repo is modified at setup time whether or not you keep the changes. A config.sample.yaml sits at the repository root as the reference for what that configuration can hold, and foreman init --force re-creates the config and reinstalls the skills and agents when you want to start from the documented defaults again. The trade is deliberate in the other direction too: because state is in your commits, it shows up in review, diffs and history alongside your own code.
Parallel workers are separated by worktree and by declared footprint
Concurrency is handled with git worktrees rather than with a queue of branches. Each parallel worker runs in its own worktree, so two workers cannot write to the same checkout at the same time. The finer control is the footprint gate: a worker declares a set of files or paths it touches, and that declared set is what keeps two jobs from colliding. The isolation is therefore declared rather than discovered, which makes the quality of your feature breakdown the limiting factor. If two features overlap in the files they claim, worktrees stop the two runs from corrupting each other's working state but the merge is still yours to resolve. The dashboard's issue board shows where that sits in practice, with columns for queued, in progress, done and merged items, and separate screens exist for watching the workers and for anything that needs your attention. For a tool whose pitch is trusting agents to merge, that division of labour is the point: the orchestrator guarantees no simultaneous writes, and the human still owns the merge.
The guardrails are the orchestrator's, not the agent's
This is the section the documentation puts first, and the reason is stated plainly: agents are boxed in by the orchestrator rather than by their own good behaviour. Three mechanisms do the boxing. Per-run caps bound turns, cost and time for a single run. A daily cost ceiling exists with a hard stop behind it, so spending past the day's limit ends the run instead of merely warning. And a PreToolUse deny hook blocks workers from writing their own verification, which closes the most tempting shortcut in agentic development, namely grading your own homework by editing the test that judges the work. Each feature is attributed to Foreman rather than to a prompt asking nicely, which is the difference between a policy and an instruction. The daily ceiling in particular is the one to set deliberately before your first real run, since it is the setting that decides how expensive a bad afternoon is allowed to become.
Retro drafts patches, and bench decides whether they land
Runs feed a small improvement loop rather than a pile of logs. Every run is outcome labelled, so the system knows which attempts produced working results and which did not. foreman retro reads that history and clusters recurring failures, turning them into gated drafts of skill or prompt patches rather than applying changes on the spot. Nothing lands until foreman bench replays the eval set and the patch passes, and bench reports three deltas rather than a single verdict: success rate, cost and turns. Requiring the turn and cost deltas matters as much as the success rate, since a patch that fixes a failure by tripling the number of agent turns has moved the cost problem rather than solved it. The skills being patched are the vendored ones that init installed into .claude/skills, and the dashboard tracks their versions, which is how you can see a fix land in the same place you started from. Two more commands round out the loop: foreman build resumes or continues the autonomous build of a feature, and foreman status reports vendored skill and agent status along with the features it knows about for the repository.
Two runtime dependencies and a console entry point
The packaging is unusually small. The project is named foreman-orchestrator on PyPI, version 0.6.0, built with hatchling, and its only runtime dependencies are textual at 0.79 or newer and pyyaml at 6.0 or newer. Everything else is optional: the dev extra carries pytest, pytest-asyncio, pytest-cov, pytest-textual-snapshot and textual-dev, so the terminal interface has snapshot testing in the toolchain. The console script maps the foreman command to foreman.cli:main, which is why one installed command covers every mode. The wheel is built to ship data as well as code, with the vendored skills, fixtures, hook assets and agents declared as artifacts, so the files the orchestrator depends on travel inside the package rather than being looked up from a checkout. Configuration is YAML given the pyyaml dependency, and pytest runs with asyncio mode set to auto. Platform support is the narrower constraint: Linux or WSL2, developed and tested on Ubuntu under WSL2. The badges at the top of the documentation point at a CI workflow, a coverage service, the PyPI listing and the license file, so continuous integration and coverage are wired up even though only two libraries are needed to run the thing.
The repository is as much process as code
The root of this project documents itself. AGENTS.md, DECISIONS.md, CHANGELOG.md, CONTRIBUTING.md, a config.sample.yaml, a docs directory, a tests directory and a validation directory all sit at the top level, alongside src, dogfood, issues, examples and a uv.lock. examples contains a hello project, which is the smallest thing worth pointing the orchestrator at before a real repository. On licensing there is a small inconsistency worth naming rather than smoothing over. The package metadata declares MIT, both as license text and as an OSI approved MIT classifier, and there are LICENSE and NOTICE files at the root, but automated license detection on the repository reports no assertion, so the authoritative answer is the LICENSE file plus the NOTICE rather than a badge. The only release is v0.6.0 on 2026-06-21, and the last push landed on 2026-06-29, with the repository not archived. The README is cut off partway through the phase table, so the later phases and the configuration section are worth reading from the repository itself. Two root directories hint at how the project is exercised: dogfood suggests running the orchestrator on its own repository, and validation suggests a separate place where runs are checked, which is consistent with a tool whose main claim is that its output can be verified.
Editorial conclusion
It fits someone already running claude in headless mode who lost an afternoon to an agent that edited past review and then approved its own work, because the hash-sealed approval and the deny hook target exactly that failure. It does not fit macOS or bare Linux users, since the stated support is Linux or WSL2 tested on Ubuntu, and it does not fit anyone who wants an agent library rather than a supervised workflow. Before you point it at a repository, read the config sample, confirm the daily cost ceiling you intend, and check that the vendored skills it writes into .claude/skills/ are ones you want in your repo.
Frequently asked questions
What does foreman init do to my repository?
It scaffolds a .foreman directory and installs the foreman prefixed skills into .claude/skills. foreman init --force re-creates the config and reinstalls the skills and agents.
Can I try foreman without spending any tokens?
Yes. foreman demo runs non-interactively and foreman --demo launches the full TUI, both against a throwaway sample repository using a mocked agent backend with canned stream-json, at zero token cost.
How does foreman stop an agent from approving its own work?
A PreToolUse deny hook blocks workers from writing their own verification, alongside per-run caps on turns, cost and time and a daily cost ceiling with a hard stop.
What does foreman require before it will run?
Python 3.11 or newer, git, the claude CLI installed and authenticated as confirmed by claude --version, and Linux or WSL2, since it was developed and tested on Ubuntu under WSL2.
What does foreman bench check before a patch lands?
It replays the eval set and reports success rate, cost and turn deltas. Patches drafted by foreman retro have to pass it before they can land.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/visionforge-ou-foreman)