Model or dataset
HangYu8123/HarnessFlow avatar
HangYu8123/HarnessFlow

HarnessFlow: a Markdown workflow pack that stops Claude Code, Codex CLI and Copilot from one-shotting your changes

Harness coding workflow for codex, claude, github copilot

462 stars25 forksHTMLLicense varies

At a glance

What is it?
HarnessFlow is a portable instruction pack, not a runtime: you copy Markdown workflow files into a repository and your AI coding assistant picks up a plan, self-challenge, QA and memory pipeline. It suits teams already using the three supported assistants and willing to fill in a request template by hand.
Who is it for?
Adopt HarnessFlow if you already drive Claude Code CLI, Codex CLI or VS Code with Copilot and you want every change to pass through analysis, a challenge pass and QA before it lands. Skip it if you need a runtime, a CI gate or a licence you can read: there is no npm install, no build step, and the repository states no licence.
Can I use it commercially?
Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
Is it still maintained?
Yes. The repository last received commits 28 days ago.
What is it written in?
Mainly HTML, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 10, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem HarnessFlow targets: assistants that go straight from prompt to diff

The README frames the failure mode plainly. A bare prompt goes to a diff with no planning stage, which means the assistant commits to an interpretation of your request before it has read enough of the codebase to know whether that interpretation is right. HarnessFlow's answer is to insert a pipeline between the request and the edit: a filled-in template names the workflow you want, analyst subagents investigate first, a challenge pass attacks the plan, and only then does code change.

The intended user is someone already running Claude Code CLI, Codex CLI or VS Code with GitHub Copilot, because those tools read instructions from the repository itself. The README also lists Aider and "other AI coding assistants", but with an explicit caveat: those follow a workflow file manually, with no subagent orchestration. That distinction matters more than it looks. The multi-agent part of HarnessFlow depends on the host tool having subagents at all, so on Aider you get the written process and lose the parallel investigation.

There is no runtime, no npm install and no build step. Everything is Markdown plus a few shell and Python helpers. That is the whole adoption surface, and it is also the whole enforcement surface: nothing in the pack can make your assistant follow the workflow.

How the pipeline works: template, parallel analysts, Devil's Advocate, QA, repo memory

The data flow described in the README runs in one direction. A request starts only when you fill in a template from request_template/, which names one of nine request types and loads the matching workflow file. Nothing fires on a bare prompt. That is a deliberate design choice: the workflow is opt-in per request, so you decide when the heavy process runs and when a quick edit does not need it.

Once a workflow file is loaded, three analyst subagents read the codebase from different angles. The README names them Focus, Broad and Free. A Senior Engineer agent then synthesizes their findings into a single plan. A Devil's Advocate pass stress-tests that plan for regressions and bad assumptions before any code is written. After implementation, a QA Engineer agent checks the result, and an opt-in approval gate lets you sign off on the plan before work begins.

Results are written to repo_info/, which is the persistent memory layer. Later requests read that directory instead of re-deriving context about the repository. The Initialize request type exists precisely for this: first-time setup bootstraps repo memory, and re-initialization validates the existing memory against the code and diff-updates it. That is the mechanism most likely to rot if you skip it, because stale memory is worse than no memory when an analyst agent trusts it.

Each of the nine request types ships in three variants. general is described as thorough, fast as token-efficient, and skill as community-skill-backed. The workflows live under workflow/ as shared sets used by all platforms, with a matching fill-in prompt in request_template/.

Installing HarnessFlow and running a first initialize request

The repository root carries setup.sh and cli_setup.sh, and the README's platform table says the entry point for each tool is generated during setup. For Claude Code CLI that is a root CLAUDE.md, for Codex CLI a root AGENTS.md, and for VS Code with Copilot .github/copilot-instructions.md. The repository already contains CLAUDE.md, AGENTS.md and copilot-instructions.md at the top level, so the generated files have visible counterparts in the tree.

The README does not print the exact arguments setup.sh accepts, so the only command traceable to the repository is the plain invocation of the script from the repository root. Run it inside the repository you want the assistant to work in, not in a separate checkout, because the whole point is that the assistant reads the instructions from the repository it is editing.

bash
bash setup.sh

After setup, confirm the entry point for your tool exists at the root. For Codex CLI, that is AGENTS.md; for Claude Code CLI, CLAUDE.md; for Copilot in VS Code, .github/copilot-instructions.md. If your tool is Aider, the README says to follow a workflow file manually instead.

The first request worth running is Initialize, because it creates the repo_info/ memory that later requests read. Pick the matching template from request_template/, fill it in, and hand it to your assistant. The template names the request type and the workflow file to load. The README does not publish a template body, so copy the structure from the file in request_template/ rather than from any example here. With the general mode you get the thorough analysts; with fast you get the token-efficient variant. The README positions fast as the efficiency-quality sweet spot, so it is the reasonable default for a first run. What you should see afterwards is a populated repo_info/ directory. If that directory stays empty, the workflow did not load and the template was probably not recognized.

Where HarnessFlow breaks down: no enforcement, no licence, and a benchmark you should read carefully

The most important limitation is structural. HarnessFlow is Markdown. It cannot stop an assistant from ignoring the workflow, skipping the Devil's Advocate pass, or writing code before the analysts finish. The README's own framing is that nothing fires on a bare prompt, which cuts both ways: the pack is inert unless you fill in a template every time you want the process. On a small fix, that overhead is real, and the fast mode exists because the authors know it.

Aider and "other LLMs" get a reduced product. The README states they follow any workflow file manually with no subagent orchestration, so the parallel Focus, Broad and Free analysis and the Senior Engineer synthesis are unavailable. If your team standardizes on Aider, you are adopting a document, not a pipeline.

The benchmark table deserves scrutiny rather than repetition. It reports 5-task total median lines of code across three models for four arms, with HarnessFlow-Fast at 58, 60 and 46 lines and native at 152, 91 and 192, and it states 180 independent single-shot generations at n=3 median. The README is unusually candid that ponytail is leaner on Haiku and Sonnet, and that ponytail slips to 40/45 on correctness while HarnessFlow-Fast holds 45/45. It also points to experiment_ponytail/REPORT.md for methodology and "honest caveats". Read that file before quoting the table. A five-task benchmark scored by the benchmark author's own scripts is a narrow basis for a workflow decision, and the README itself says the model is held constant per arm so the comparison isolates the harness.

The repository states no licence. The README does not mention one, and the top-level entries list no LICENSE file. If you plan to redistribute the pack inside a company or ship it in a product, that is a question to resolve with the copyright holder before you copy the files, not after.

HarnessFlow compared with ponytail and fastworkflow

The README benchmarks HarnessFlow against two other harnesses, and the difference in approach is the interesting part. ponytail is described as a minimal-code skill: its goal is to make the model write less code, and on Haiku and Sonnet it does that better than HarnessFlow-Fast, at 37 and 50 lines against 58 and 60. The cost shows up in correctness, where ponytail reaches 40/45 and occasionally ships broken code according to the README. HarnessFlow-Fast trades a little leanness for 45/45.

fastworkflow sits at the opposite end. Its validation-first style writes 1.6 to 3.4 times more code than the bare model, with 288, 305 and 313 lines against native's 152, 91 and 192, and it is the least correct arm at 31/45. That is the clearest signal in the table: adding process does not automatically improve output, and a validation-heavy harness can inflate the diff while lowering correctness.

The structural difference beyond the numbers is that HarnessFlow is a portable instruction pack rather than a framework or a skill. ponytail is a skill, fastworkflow is described as a framework, and HarnessFlow is a set of Markdown files you copy into a repository so three specific assistants can read them. If you want a single skill you can install once and forget, the other two are closer to that shape. If you want a per-request process with an audit trail in repo_info/, HarnessFlow is the one built for that.

Maintenance, upgrade cost and what the repository does not say

The last push to the default branch was on 2026-09-12, eight days before this writing, and the repository is not archived. There are no releases retrieved, so upgrades are not versioned artifacts you pin. You take the current state of main, or you take nothing.

That shapes the upgrade path. Because the pack is copied into your repository, upgrading means re-copying files and reconciling them with whatever local edits you made to the workflows or templates. The repository contains sync_agent_definitions.py and sync_gui_templates.py, which suggests the authors maintain generated artifacts across platforms, but the README does not document an upgrade procedure, and it does not document rollback. If you customize a workflow file, you own that divergence.

The maintenance cost that is documented is the memory. repo_info/ is written by completed requests and read by later ones, and the Initialize request type explicitly covers re-initialization: validate the existing memory against the code and diff-update it. A repository that changes quickly will need that pass regularly, or the analyst agents will plan against a description of the codebase that no longer matches it. Nothing in the README schedules that for you.

On licence, the honest position is that there is nothing to read. No licence identifier appears in the repository description, and no LICENSE file appears in the top-level entries. Absence of a licence is not permission. Treat redistribution and internal policy questions as unresolved until the copyright holder states terms.

Editorial conclusion

Adopt HarnessFlow if you already drive Claude Code CLI, Codex CLI or VS Code with Copilot and you want every change to pass through analysis, a challenge pass and QA before it lands. Skip it if you need a runtime, a CI gate or a licence you can read: there is no npm install, no build step, and the repository states no licence. Before you copy anything in, open request_template/, check that the nine request types cover the work you actually do, and read the benchmark caveats in experiment_ponytail/REPORT.md rather than the summary table.

Frequently asked questions

Is HarnessFlow a CI/CD tool?

No. It is a portable Markdown instruction pack with no runtime, no npm install and no build step, and it runs inside your AI coding assistant rather than in a pipeline. It has no hooks into CI systems or pull request checks.

Which AI agent harness is the best?

The README does not rank harnesses in general, and its benchmark is limited to four arms on a five-task code-generation benchmark. In that comparison, HarnessFlow-Fast is described as the only harness arm that is both leaner than the bare model and fully correct, while ponytail is leaner on two of three models and fastworkflow writes far more code.

Is HarnessFlow similar to GitHub?

No. HarnessFlow is a set of workflow and agent-instruction files you copy into a repository, and it works with assistants such as Claude Code CLI, Codex CLI and VS Code with Copilot. It is not a hosting platform and does not store or serve repositories.

Official sources

  1. HangYu8123/HarnessFlow on GitHub
  2. Issues
  3. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/hangyu8123-harnessflow.svg)](https://hysenlabs.com/projects/hangyu8123-harnessflow)