# gauntlet-loop ships one SKILL.md and a prompt you paste yourself

> gauntlet-loop packages Matt Shumer's gauntlet loop as a reusable Claude skill. It runs nothing itself: it checks the quality bar you name three ways, writes a prompt of about 150 words, stops, and hands that prompt to a fresh session.

**robonuggets/gauntlet-loop** — Turn any goal into a short prompt that makes your agent set a real quality bar, run builder and critic pairs, compare blind, and loop until it wins.

- Repository: https://github.com/robonuggets/gauntlet-loop
- Stars: 1,012 · Forks: 89
- Language: Unknown
- License: CC-BY-4.0
- Published: 2026-09-16 · Updated: 2026-09-16 · Language: en
- Canonical page: https://hysenlabs.com/projects/robonuggets-gauntlet-loop

## The whole skill is one file, and the repo ships three others

The included-files listing is short enough to read in full, and its own comment makes the point:

```
.claude/skills/gauntlet-loop/
└── SKILL.md      # the whole skill, one file
README.md
LICENSE           # CC BY 4.0
```

There is no script, no runtime, no package manifest and no test suite. The behaviour you are installing is entirely contained in a single Markdown file that your agent reads.

That listing, however, is not the whole repository. The top level holds `.claude/`, `LICENSE`, `README.md` and an `assets/` directory, and the assets directory appears nowhere in the files section. Whatever is in it is not described by the page that tells you what you are about to copy.

The licence is CC BY 4.0, stated as free to use with attribution, and the credit line splits authorship in two: the skill by Jay E at RoboNuggets, the technique by Matt Shumer.

## The skill writes a prompt of about 150 words and stops

The most important thing to understand before installing anything is that this skill does not run the loop. It produces text.

The sequence it walks you through is four steps. You give it a goal, anything from a site to an essay to a research brief. It offers two or three quality bars, each one a specific real thing rather than a category. You pick one. It then writes a single short prompt of around 150 words and stops.

That prompt is what does the work, and you paste it into a fresh session, where the agent splits the work, runs builder and critic pairs, and loops until the comparison is won. So there are two agents in play and they are not the same one: the first writes the instructions, the second obeys them.

The stated reason for the fresh session is context. A critic that has watched a builder struggle carries that struggle into its judgement, and the skill treats that contamination as a failure mode rather than a nuance.

## The bar is checked three ways before any prompt is written

A bar is the reference your work gets compared against, and the skill refuses a vague one. It checks three properties before writing anything.

Named means a specific thing rather than a category, so a particular page, post or repository instead of a phrase like award-winning design. Fetchable means the critic can screenshot it, read it, run it or open it, and the reason given is the sharpest line in the project: if the agent cannot get the reference, it hallucinates the comparison and approves everything. Comparable means both sides can sit next to each other and a judge can pick one.

That middle check is the one that carries the design. A comparison the critic cannot perform is not a weak comparison, it is a fabricated one, and a fabricated comparison always passes. The skill therefore spends its validation budget on whether the reference is reachable at all before it spends anything on how strict the comparison is.

The documented examples follow the same discipline. A running brand becomes a specific brand's live campaign page, screenshotted at desktop and mobile. A CLI becomes a named tool's implementation plus its benchmark, so taste and a number both have to win.

## The critic returns a pick, not a score out of ten

The critic is described as the part that matters, and its design is a deliberate subtraction. It is a separate agent with fresh context. It opens the actual output rather than trusting a summary. It puts your work next to the bar with the labels stripped. And it says which one is better.

The stated reason it does not return a score is the useful part: a score out of ten drifts upward every round. A rating has no fixed point, so a critic that returns 7 then 8 then 9 has produced a trend rather than a verdict, and a trend is easy to mistake for progress. A binary pick cannot drift, because the answer space has two elements.

Stripping the labels is the second half of the same idea. Without it, the critic knows which side is yours, and knowing is enough to move the result.

The exit condition follows from the comparison rather than from a budget. The loop ends when your work wins the blind comparison, or when you stop the run. Never after a fixed number of rounds.

## The install path is hardcoded to .claude/skills

Portability is claimed in one section and undercut in another. The quick start is Claude-specific in both its steps:

```bash
git clone https://github.com/robonuggets/gauntlet-loop
```

```bash
cp -r gauntlet-loop/.claude/skills/gauntlet-loop your-project/.claude/skills/
```

Then it is invoked as a slash command:

```bash
/gauntlet-loop build me a pricing page for my SaaS
```

The portability claim is about the prompt the skill writes, not about the skill's location. The text says that `/loop` and `ultracode` are Claude Code features, and that for any other agent the skill swaps those two lines for plain instructions: keep looping until the critic picks ours, and run the builders and critics as parallel subagents. The structure is identical.

So on a non-Claude agent the generated prompt adapts, while the destination directory in the copy command still says `.claude/skills/`. Nothing in the quick start says where the file goes instead. That is a one-line change for whoever installs it, and it is the first thing that will trip up a Codex or OpenCode user.

## No releases and no tags, so a copied skill tracks main

The repository publishes no GitHub releases, so there is nothing to pin and nothing to diff against. The install procedure is a clone followed by a recursive copy, which means the version you get is whatever the default branch held at the moment you ran it.

For a repository whose entire payload is one Markdown file, that is a mild problem. For a repository whose payload is a prompt that decides how your agent grades itself, it is worth thinking about: a skill that changed its exit condition or its critic rules would change your results, and a cloned copy gives you no marker for what changed.

The last push to the default branch was on 2026-08-05. That is the only date the project offers as an anchor.

There is also no published versioning scheme in the file listing, no changelog and no directory-level structure to diff. For a team that wants the behaviour pinned, the practical move is to keep the copied `SKILL.md` under your own version control and treat upstream as a suggestion.

## The repository says plainly that it is not the technique

The credit section is unusually direct about origin. The gauntlet loop technique belongs to Matt Shumer, who wrote the original prompt and named the loop while building Claude of Duty. The harsh critic, the blind comparison and the refusal to stop until the work wins are all attributed to that prompt, with a link to it. The repository's own claim is narrower: it is a skill that writes a gauntlet loop prompt for you, for any goal, so you do not have to hand-write one each time.

The related reading points at Anthropic's piece on building effective agents, which covers the evaluator pattern the loop is built on. That is the useful citation to carry forward, because it is about the pattern rather than about this packaging.

The failure list is the most practically useful part of the whole page, and all four items are about discipline rather than tooling: a vague bar, which is named as by far the most common failure, a builder judging its own work, a soft critic given a scoring job instead of a binary one, and a fixed round count.

## Conclusion

gauntlet-loop is worth copying into a repository when your agent keeps stopping at good enough and you have a real reference to compare against, since the skill refuses a bar it cannot fetch and that refusal is the whole value. It is not a tool for projects with no external standard to aim at, and it is not something you can version, because the install is a clone plus a copy and the repository publishes no releases. Before you adopt it, check three things: that the bar you want is a named artifact your critic can actually open, that you can accept a critic with fresh context telling you your work lost, and that the loop has a stop condition you are willing to enforce yourself, since the design explicitly refuses a fixed round count. The last push to main was on 2026-08-05.

## FAQ

### What is gauntlet loop AI?

It is a prompting pattern, and this repository packages it as a reusable skill. A goal is turned into one short prompt of around 150 words, and that prompt tells an agent to split the work, run a builder and a separate harsh critic on each piece, compare the result against a real reference with the labels stripped, and keep looping until the work wins.

### What is gauntlet loop in this repository?

It is a skill that writes a gauntlet loop prompt for you rather than the technique itself. The technique is credited to Matt Shumer and his original prompt, and the repository's stated role is to save you hand-writing a new prompt each time, for any goal.

### How do I install the gauntlet-loop skill?

Clone `https://github.com/robonuggets/gauntlet-loop`, then copy the skill folder into your own project with `cp -r gauntlet-loop/.claude/skills/gauntlet-loop your-project/.claude/skills/`. The skill itself is a single file, `SKILL.md`, and there is nothing else to build or install.

### Does gauntlet-loop work outside Claude Code?

The generated prompt does. `/loop` and `ultracode` are Claude Code features, so for any other agent the skill swaps those two lines for plain instructions: keep looping until the critic picks ours, and run the builders and critics as parallel subagents. The structure is identical, though the copy command in the quick start still points at a `.claude/skills/` directory.

### How does gauntlet-loop decide that it is finished?

The loop exits when your work wins the blind comparison, or when you stop the run. It never ends after a fixed number of rounds, and a fixed round count is listed among the four things that break the method.

## Sources

- [Issues](https://github.com/robonuggets/gauntlet-loop/issues)
- [License: CC-BY-4.0](https://github.com/robonuggets/gauntlet-loop/blob/main/LICENSE)
- [README](https://github.com/robonuggets/gauntlet-loop/blob/main/README.md)
- [robonuggets/gauntlet-loop on GitHub](https://github.com/robonuggets/gauntlet-loop)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/robonuggets-gauntlet-loop
