loop-engineering: a pattern library and CLI for running coding agents on a loop
Practical patterns, starters & CLI tools for loop engineering with AI coding agents. Design systems that prompt and orchestrate agents (inspired by Addy Osmani and Boris Cherny). Includes loop-audit, loop-init, loop-cost.
At a glance
- What is it?
- The repository packages eight operating patterns, a scaffold CLI and an audit score for teams that want agents to find work instead of waiting for prompts. It is a design toolkit, not an autonomous refactor button.
- Who is it for?
- Adopt it if you already run Claude Code, Codex, Grok or a similar agent against a real repository and you want a repeatable way to decide what the agent does each day, how it verifies its own work, and what it is allowed to change. Do not adopt it if you want a tool that rewrites a module for you; the README states plainly that this is not a rewrite button, and the patterns are report-first for a reason.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 2 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem loop-engineering addresses: prompt fatigue around a live codebase
Most teams using coding agents operate the same way. A person decides what to do next, types it, reads the diff, and decides again. That works for a single feature. It scales badly across issues, CI failures, dependency bumps and stale pull requests, because the human is the scheduler. loop-engineering takes the position that the schedule itself should be designed and stored in the repository, so the agent has a defined job when it wakes up.
The README frames this as designing a system that discovers work, hands it to agents, verifies results, and persists state, instead of typing the next prompt yourself. The tagline is blunt about the intended audience: stop prompting, design the loop, get a score. It is aimed at engineers who already run agents against a repository and want the operating discipline written down, not at people looking for a first introduction to AI coding.
The repository is explicit about what it is not. The README says it is a pattern library for operating agents around a codebase, and adds that it is not a rewrite the module button. That distinction matters when you evaluate it, because the value is in the patterns and the readiness score rather than in code generation.
How the loop is structured: patterns, levels, and a state file
The core abstraction is a named pattern with a cadence. The registry lists eight: Daily Triage at one day to two hours, Thin loop at event plus one day, PR Babysitter at five to fifteen minutes, CI Sweeper at five to fifteen minutes, Dependency Sweeper at six hours to one day, Changelog Drafter at one day or per tag, Post-Merge Cleanup at one day to six hours, and Issue Triage at two hours to one day. Each pattern carries a week-one level and a cost band, from very low for the thin loop to very high for the CI Sweeper.
The maturity model has three levels. L1 is report-only, L2 is assisted, and L3 is unattended. The README states that week one is report-only, and that you move from L1 to L2 to L3 only after the verifier has been right for a week. That sequencing is the most useful idea in the repository, because it turns trust in the agent into something you earn per loop rather than assume.
State persistence is handled through a STATE.md file at the repository root. The README notes that the Loop Ready score now weights recent runs harder than files on disk, and gives a concrete example: a 30-day-old STATE.md is not L3. In other words, a stale state file does not count as evidence that the loop is running. The repository also ships a thin GitHub Action starter under starters/thin-loop/ which the README describes as requiring no STATE.md, so state persistence is optional for the lightest pattern.
Installing loop-engineering and running a first daily triage loop
The CLI is published as an npm package and is meant to be invoked through npx, so there is no global install step. The first command scaffolds a pattern into the current directory. The README gives `--tool claude` in its example and states that `--tool` defaults to `claude` if you omit it, and that you can swap in `grok`, `codex`, or `opencode`.
npx @cobusgreyling/loop init . --pattern daily-triage --tool claudeAfter that, the doctor command inspects the scaffolded loop. The README pairs the two commands in its getting-started block, and the Quickstart links to a walkthrough script, scripts/empty-to-state-demo.sh, that takes an empty repository through the first STATE.md.
npx @cobusgreyling/loop doctor .The third command in the same block estimates what the pattern will cost at a given level. This is the step most people skip, and it is the one that tells you whether a five-minute cadence is realistic for your budget.
npx @cobusgreyling/loop cost --pattern daily-triage --level L1The unified front door exposes five subcommands according to the README: init, doctor, status, audit, and cost. Older packages such as loop-init and loop-audit stay supported. What you should see after running init is a scaffolded loop in your working directory plus a STATE.md, and the Quickstart is the document to follow for the walkthrough. The README does not document a rollback command for a scaffolded loop, so plan on removing generated files by hand if you change your mind.
Loop Ready scoring and what it actually measures
The audit command produces a Loop Ready score, and the README's demo asset is a GIF of that score climbing. The scoring rule that matters is the recency weighting described above. A loop that ran last week scores differently from one whose only evidence is a month-old state file, even if both have the same files on disk.
That is a defensible design choice and also a source of confusion. A team that runs a loop manually once a month will see a lower score than a team running a daily triage cadence, because the score is measuring operating behaviour rather than repository contents. If you adopt the audit as a gate in CI, be aware that you are gating on activity, not on configuration.
The repository also carries a gate.yaml at the root and a loop-gate tool with its own test script in package.json, which suggests the score can be wired into a pass or fail check. The README does not spell out the gate's exact thresholds, so read gate.yaml directly rather than assuming a number.
Where loop-engineering is the wrong tool
The README is unusually candid about failure. It states that loop engineering amplifies judgment, that token costs can explode, that unattended loops make unattended mistakes, and that you should read what the loop ships. Those are not hedges added for legal cover; they describe the actual failure mode of running an agent on a timer against a repository.
The cost bands in the pattern table are the practical constraint. PR Babysitter is listed as high cost and CI Sweeper as very high, because both poll on a five to fifteen minute cadence. A solo developer on a personal project will not get value from either. The same applies to Dependency Sweeper at medium cost with a six-hour to one-day cadence and a patch-only week-one level.
If your problem is a single large refactor, this repository is the wrong shape. The README points that use case at docs/refactor.md and describes shipping a feature or refactor as a sequence of todos into small pull requests, which is a different workflow from the continuous patterns. And if you have no verifier you trust, the L1 to L2 to L3 ladder has nothing to climb on. The README's own rule, that you promote only after the verifier has been right for a week, means a team without a reliable check stays at report-only indefinitely, which is the correct outcome but not the one people expect when they install a loop tool.
How this differs from prompt engineering and from harness or graph approaches
Prompt engineering optimises a single interaction: the wording you send to a model. loop-engineering optimises the schedule and the verification around many interactions. The README's own framing, design the system that prompts your agents, is the cleanest statement of the difference. You are not writing a better prompt; you are deciding which prompt gets sent, when, by what, and what happens to the result.
Against harness engineering, the distinction is where the machinery lives. A harness is the runtime that executes an agent's tool calls. loop-engineering sits above that and describes the cadence, the level of autonomy and the state file, and it is agnostic about which agent runtime you use, which is why the examples directory carries separate folders for Claude Code, Codex, Cursor, GitHub Actions, Grok, Hermes, MCP, OpenClaw, Opencode and Windsurf. A graph-based approach models the task as nodes and edges; the patterns here are closer to cron jobs with a verification step and a promotion ladder.
The repository also lists companion projects for later stages, including memory-engineering, harness-foundry, outerloop, fleet-engineering and goal-engineering, with an explicit instruction not to add them until a loop has actually run. That ordering is a real opinion: adopt the smallest loop first.
Licence, maintenance and the cost of keeping a loop running
The repository is MIT licensed, which permits commercial use and modification subject to the licence terms. The npm package is published under the @cobusgreyling scope. Nothing in the README suggests a separate commercial tier or a licence key, and the companion repositories are linked as open projects rather than paid add-ons. This is not legal advice; read the LICENSE file for the actual terms.
The last push to the default branch was on 2026-09-09, and the most recent release listed is v1.6.0 from 2026-07-20, titled Foundry funnel + loop-gate. v1.5.0, the Community Tools Drop, landed on 2026-06-30. The repository is not archived.
Upgrade cost is mostly in the tooling rather than the patterns. The root package.json defines separate test scripts per tool, including test:loop-audit, test:loop-init, test:loop-cost, test:loop-gate, test:mcp-server, test:loop-sandbox and others, plus a test:tools aggregate and a build:tools chain that builds each tool in sequence. If you vendor any of these tools, that is the surface you inherit. The README also states that older packages such as loop-init and loop-audit remain supported alongside the unified loop command, so a migration is available but not forced.
The other recurring cost is tokens. The README warns that token costs can explode, and the cost subcommand exists precisely so you can price a pattern and level before enabling it. Budgeting is a first-class part of the workflow here, not an afterthought.
Editorial conclusion
Adopt it if you already run Claude Code, Codex, Grok or a similar agent against a real repository and you want a repeatable way to decide what the agent does each day, how it verifies its own work, and what it is allowed to change. Do not adopt it if you want a tool that rewrites a module for you; the README states plainly that this is not a rewrite button, and the patterns are report-first for a reason. Before you commit, verify three things against the repository: which pattern matches your actual cadence, whether your verifier can be trusted for a week at L1 before you move to L2, and what the cost command reports for your chosen pattern and level. The gate file at the repository root, gate.yaml, and the loop budget file, loop-budget.md, are the concrete artefacts to read before enabling anything unattended.
Frequently asked questions
What is loop engineering in AI?
It is the practice of designing the system that discovers work, hands it to AI coding agents, verifies the results and persists state, rather than typing each prompt yourself. The repository frames it as designing the loop instead of prompting.
What are the key differences between loop engineering and prompt engineering?
Prompt engineering optimises a single interaction with a model. loop-engineering optimises the cadence, autonomy level and verification around many interactions, with the agent's job defined in the repository rather than decided per prompt.
How do I set up loop engineering?
Run the scaffold command for a pattern and tool, then run the doctor command against the directory, then check the cost for the pattern and level. The README's example uses daily-triage with claude, and week one is report-only.
How do I use loop engineering in Claude Code?
The README shows init with --tool claude, which is also the default when the flag is omitted. The examples directory has a Claude Code folder that includes a plugin page, and the Quickstart is the walkthrough to follow.
Can you provide an example of loop engineering?
Daily Triage is the entry pattern: a one-day to two-hour cadence, an L1 report-only first week, and a low cost band. It targets repository health such as issues, CI and dependencies.
Who is the CEO of loop AI?
The repository does not address this. It documents patterns, CLI tools and a readiness score for running coding agents in loops, and names no company leadership.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/cobusgreyling-loop-engineering)