# architect-loop: an autonomous build loop for Codex and Claude

> architect-loop turns a goal into a shipping pull request using a fresh strategist, isolated builder worktrees and frozen checks. The README is explicit about the sandbox being weak, which is the first thing to weigh.

**DanMcInerney/architect-loop** — Super optimized /goal loop. Massive token savings and higher quality. Smart model designs and reviews, cheaper model builds.. 

- Repository: https://github.com/DanMcInerney/architect-loop
- Stars: 626 · Forks: 53
- Language: Python
- License: MIT
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/danmcinerney-architect-loop

## The problem architect-loop targets: long agent runs that lose the plot

A coding agent asked to build a multi-hour feature tends to fail in a predictable way. Context fills with half-finished edits, the model starts grading its own work, and the final answer is a summary rather than a diff. architect-loop is built around that failure. The README describes it as "an autonomous software factory for Codex and Claude: one orchestrator, fresh-context strategist and builder agents, deterministic gates, and a single shipping PR or local finish record." The intended user is someone who already drives Codex or Claude as a coding agent and wants to hand over a feature-sized request rather than a single edit.

The unit of work is the issue, not the conversation. A configurable strategist subagent handles full-lane design and review, and fresh builders each implement exactly one issue in an isolated worktree. Because builders are fresh, they do not inherit the strategist's reasoning or the orchestrator's history. That is the whole point: no agent grades its own work in the same context window that produced the work.

## How the /architect lane moves from goal to frozen checks

The README lays out the flow as a sequence of stages. Intake feeds the goal straight to a fresh strategist that drafts the spec. Open questions do not block: they become timed rulings that auto-default into recorded assumptions. There is no human approval gate anywhere, and the README says the human steers by editing the spec, commenting rulings, or creating `docs/STOP`.

A second, fresh harden strategist attacks the draft, folds surviving findings into a revised spec, decomposes it into file-disjoint vertical-slice issues, drafts frozen checks, and stress-tests its own decomposition. The orchestrator then publishes the spec and issues to the tracker and owns the freeze commit. Checks freeze in git before any builder exists, which is the mechanism that makes later grading independent of later code.

Run identity is pinned by `docs/runs/<run>/manifest.md`, which records the tracking issue, tracker mode, factory branch, and spec path. Status commands take a run slug and no longer discover the current run by scanning issues. GitHub and markdown tracker modes share the same state machine, and markdown issues live in `docs/issues/<run>/` with no GitHub remote required.

Builders run in fresh worktrees. They must state execution conflicts before coding, must not commit, and must end with one raw `STATUS:` line. The README is blunt that exit truth outranks terminal-looking report text. Wrapped CLI jobs write `job.meta.json`, `job.heartbeat`, `job.exit.json`, `events.jsonl`, and `stderr.log`, and postflight refuses to merge CLI job work without that wrapper exit truth.

## Installing architect-loop and running a first bounded change

The README gives a three-line install. The installer copies the same `skills/` tree to Claude and Codex skill locations, so one install serves both agents.

```bash
git clone https://github.com/DanMcInerney/architect-loop
cd architect-loop && ./install.sh        # Windows: .\install.ps1
npm i -g @openai/codex@latest            # optional builder backend
```

On Windows the equivalent script is `install.ps1`. If you would rather not touch your user profile, the README states that `./install.sh --project` or `.\install.ps1 -Project` installs into the current repo instead. The npm line is marked optional and only matters if you want Codex as the builder backend.

Three commands are exposed:

```text
/architect-research <topic>
/architect <hours+ feature/product request>
/architect-fast <small change>
```

For a first real use, `/architect-fast` is the sensible entry. The README describes it as shipping a small, bounded change through the light lane, capped at at most three builder issues and roughly 400 changed lines. It skips strategist subagents, adversarial review, frozen checks, the per-issue check-runner, and the watchdog script. Issue-body acceptance criteria, builder-run tests, and the closing builder review carry the gate instead. If the work exceeds the size ceiling, the lane stops and recommends `/architect`.

`/architect-research` is separate: a scout maps brainstorm-scale topics before lanes are designed, researchers gather in parallel under hard budgets, and the draft is marked SUPPORTED, THIN, or EMPTY, with thin or empty sections triggering up to two targeted gap rounds.

## The weak sandbox is the real limitation, and the README says so

Most agent frameworks describe their safety model in confident terms. This one does the opposite. The README states that the OS sandbox is deliberately weak (`danger-full-access`) and that builders may use the network and the wider filesystem. The boundaries it actually relies on are the postflight touch-set audit, the never-commit ban, and read-only frozen checks.

That is a real trade-off, not a marketing line. Builders can reach the network and files outside the worktree, and the compensating controls are audits that run after the fact. A builder that writes outside its expected touch set is caught at postflight, which means the detection is retrospective. If your threat model requires preventing an agent from touching the network at all, this design is the wrong fit regardless of how good the gates are.

The second limitation is operational. Because postflight refuses to merge CLI job work without wrapper exit truth, an unwrapped job cannot merge even if its code is correct. The README notes the watchdog detects legacy unwrapped jobs among the stall and failure patterns it reports. If you invoke jobs outside `run-job.ps1|.sh`, you get work that the pipeline will not accept.

A third constraint is that the watchdog never acts. It reports typed evidence only: stalls, repeated commands, blocked tool calls, orphaned wrappers, failed exits, and legacy unwrapped jobs. It never kills, nudges, or grades. Stuck jobs are reaped manually with `kill-job.ps1|.sh` and respawned from durable issue context. The orchestrator still decides. Anyone expecting a self-healing loop should read that paragraph twice.

## How architect-loop differs from a plain agent loop

The obvious alternative is running Codex or Claude directly and letting it iterate on a task until it stops. The difference is where judgment sits. In a plain loop, the same context window that wrote the code decides whether the code is good, and the stopping condition is usually the model's own claim of completion. In architect-loop, the closing `final-review` is read-only and returns `REVIEW: GREEN` or a review spec plus fix issues. Findings become fix issues and draft checks for a fix wave, and zero findings short-circuit to integrate. The reviewer cannot edit; it can only file work.

The second difference is that the gate is frozen before the work starts. Checks freeze in git before any builder exists, and the check-runner grades frozen RUN items with fixed expectations. A builder cannot weaken the test it will be judged against, because the test was committed by the orchestrator first. That is a structural answer to reward hacking rather than a prompt-level one.

The third is scope discipline. Builders get one issue each in an isolated worktree, and issues are decomposed to be file-disjoint. That avoids the merge conflicts you get when several agents edit the same file in parallel, at the cost of requiring the strategist to actually produce a clean decomposition. The README says the harden strategist stress-tests its own decomposition, which suggests the project treats that as a known weak point.

A note on lineage: the README states that orchflows now replaces this hardcoded workflow with a library that can build any kind of workflow like this with only 2 skills. The author is pointing at a successor. That is worth reading before committing to this repository long term.

## Maintenance, licence, and what a run leaves behind

The repository is not archived and the last push was on 2026-09-13, four days before this writing. That is a current codebase, but the README's own note about orchflows superseding the hardcoded workflow means the upgrade path may point at a different repository rather than a new tag here. No releases were retrieved, so there is no version number to pin against. You are tracking `main`.

The licence is MIT, listed at the top level as `LICENSE`. MIT permits commercial use, modification, and redistribution with the licence and copyright notice retained. That is the general shape of the licence, not legal advice; read the file for the operative text.

Upgrade cost is mostly re-running the installer, since it copies the `skills/` tree to Claude and Codex skill locations. The stage-skill library listed in the README is `codebase-design`, `to-spec`, `adversarial-review`, `to-issues`, `frozen-checks`, `tdd`, `final-review`, and `integrate`. Any local edits to those skill files would be overwritten by a reinstall, so keep customizations outside that tree.

Run artifacts are durable and worth planning for. `docs/runs/<run>/manifest.md` pins run identity, `docs/issues/<run>/` holds markdown issues in that tracker mode, and wrapped jobs write `job.meta.json`, `job.heartbeat`, `job.exit.json`, `events.jsonl`, and `stderr.log`. Postflight can defer locked worktree cleanup, and `sweep-deferred.ps1|.sh <run>` clears run-scoped debris before final close. On a busy machine those deferred worktrees accumulate until swept.

## What to check before you point architect-loop at a real repository

Three things decide whether the first run succeeds. First, confirm the installer landed where your agent reads skills from; the README notes both Claude and Codex locations are written, and `--project` changes that target. Second, decide your tracker mode up front. GitHub and markdown share the same state machine, but markdown issues live under `docs/issues/<run>/` and need no GitHub remote, so the choice affects whether the run can post a PR at all. Third, make sure your jobs go through `run-job.ps1|.sh`, because postflight will not merge CLI job work without `job.exit.json`.

The closing sequence is worth understanding before you start, since it determines what you get. The orchestrator runs the closing test pass over builder-built suites plus every frozen RUN item, then feeds the raw output into a single read-only final review. Findings become fix issues and draft checks for a fix wave, and zero findings short-circuit to integrate. `integrate` is a builder-model stage whose first step is the docs pass; it consumes the change-context digest, updates product docs, verifies the closing state, sweeps deferred cleanup, and prepares the PR or markdown finish record.

## Conclusion

Adopt architect-loop if you already run Codex or Claude as a coding agent and want multi-hour feature work decomposed into file-disjoint issues with frozen checks and a shipping PR, and you accept the README's own statement that the OS sandbox is deliberately weak. Do not adopt it for small edits, since /architect-fast exists precisely because the full lane is overkill, and do not adopt it if you need a human approval gate, because the README says there is none anywhere. Verify first that install.sh copies the skills tree to the location your agent reads, that your tracker mode (GitHub or markdown) matches your remote setup, and that postflight will accept your jobs, since wrapped CLI jobs must produce job.exit.json before their work can merge.

## FAQ

### What is architect-loop and what does it do?

It is an autonomous software factory for Codex and Claude, built around one orchestrator, fresh-context strategist and builder agents, deterministic gates, and a single shipping PR or local finish record. The README states there are no approval gates, with timed rulings throughout.

### How does the architect-loop loop actually work?

Intake sends the goal to a fresh strategist that drafts a spec, a second harden strategist attacks it and decomposes it into file-disjoint issues, and the orchestrator freezes checks in git before any builder exists. Builders then implement one issue each in isolated worktrees and report raw evidence.

### What are the different lanes in architect-loop?

The README lists three: /architect-research for answer-first reports, /architect for the full multi-hour factory, and /architect-fast for small bounded changes capped at three builder issues and roughly 400 changed lines. If /architect-fast work exceeds that ceiling, it stops and recommends /architect.

### What is a loop in engineering, in the context of architect-loop?

In this project the loop is the full cycle from intake to shipping PR: a strategist drafts a spec, a harden strategist decomposes it, builders implement one issue each in isolated worktrees, and a read-only final review returns REVIEW: GREEN or files fix issues.

## Sources

- [DanMcInerney/architect-loop on GitHub](https://github.com/DanMcInerney/architect-loop)
- [Issues](https://github.com/DanMcInerney/architect-loop/issues)
- [License: MIT](https://github.com/DanMcInerney/architect-loop/blob/main/LICENSE)
- [README](https://github.com/DanMcInerney/architect-loop/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/danmcinerney-architect-loop
