# ralph-playbook: an agent loop designed backwards from a context budget

> The Ralph Playbook is a method document with no code in it, and its central claim is arithmetic rather than mystical: an advertised context window of two hundred thousand tokens is really about a hundred and seventy six thousand, the useful zone is forty to sixty percent of that, so the loop runs exactly one task per iteration to stay inside it. Everything else in the guide follows from that number.

**ClaytonFarr/ralph-playbook** — A comprehensive guide to running autonomous AI coding loops using Geoff Huntley's Ralph methodology. View as formatted guide below 👇

- Repository: https://github.com/ClaytonFarr/ralph-playbook
- Website: https://claytonfarr.github.io/ralph-playbook/
- Stars: 1,034 · Forks: 265
- Language: HTML
- License: MIT
- Published: 2026-09-30 · Updated: 2026-09-30 · Language: en
- Canonical page: https://hysenlabs.com/projects/claytonfarr-ralph-playbook

## A document that opens by describing how it was wrong

The introduction of this guide is unusual, and it is the reason to trust the rest of it.

It opens by dating the moment the method it describes went viral, in December 2025. It then admits that the author spent the summer before that unable to make the method click for them, and that two other people's overviews of it helped right up until the method's originator publicly disagreed with them. The author embedded a screenshot of that disagreement in the repository, in a directory called references, alongside the diagram the guide is built around.

Then comes the stated method of the work: watch the recent videos, read the original post, and try to work out for himself what actually works, rather than what the summaries said. The word used for reading the source material is a four-letter internet abbreviation for reading the manual, and the self-description of the result is a guide organised so it can be put into practice, hopefully without neutering the approach in the process.

That is a small thing and it changes how you should read a technical opinion piece. Most guides of this kind open with a thesis and cite supporting sources chosen after the fact. This one opens with a correction, cites the person who made it, and then goes back to primary material. The rest of the document is correspondingly specific: it contains numbers, file names, a step count, a vocabulary with definitions, and a test with pass and fail examples. It is arguable with, which is the highest compliment you can pay a method document.

The provenance discipline continues in the repository itself. There is a notice file next to the licence, which is not where a document like this normally has one, and the guide credits three people by name with links to their original material. The methodology is not the author's; the synthesis is. Saying so in the first paragraph rather than in a footnote is the whole point.

## The context arithmetic is the part you can reuse tomorrow

The key principles section opens with a claim that everything else in the guide depends on: context is everything. And then it does something unusual for this genre, which is to put numbers on it.

The arithmetic runs like this. When a model advertises two hundred thousand tokens or more of context, the figure you can actually use is around a hundred and seventy six thousand. Within that, somewhere between forty and sixty percent is described as the smart zone, the band where output quality is good. And the design conclusion follows: if your tasks are tight and you run exactly one task per loop iteration, you land at full utilisation of the smart zone.

That last step is the interesting one, because it converts a quality heuristic into a scheduling policy. If quality degrades as a context window fills, then the number of tasks per iteration is not a matter of throughput preference, it is the primary control variable. Running five tasks per iteration to get more work done per cycle means running every one of them in the bottom of the window. The guide's position is that you should not do that, and it says so by doing the division rather than by asserting a principle.

Two more derived practices follow from the same numbers, and they are about where work happens. The main context is treated as a scheduler rather than a worker: expensive work is pushed out to subagents instead of being done in the main thread, because the main thread is the one whose budget you are protecting. And subagents are described as a memory extension, with each one given roughly a hundred and fifty six kilobytes that is discarded when it finishes.

That figure is worth a second look, because it is nearly the whole usable window. One subagent at that size and the main context has almost nothing left. Which means the guidance is not really about running many subagents in parallel, it is about running one at a time and throwing it away, and the guide's own summary of the subagent role, memory extension, is the accurate description of that rather than parallelism.

A third practice is about determinism rather than budget: verbose inputs degrade it. The recommendation is fewer parts, less configuration, shorter content, and markdown rather than structured data formats for defining and tracking work, on the grounds of token efficiency. This is the least discussed point in the guide and possibly the most important, because a loop that behaves differently between runs is a loop you cannot debug, and reducing input surface area is the cheapest way to reduce that.

## One loop, two prompts, and a planning mode that usually stops after two turns

The structural claim in the guide is that this is not a loop that writes code. It is a funnel with three phases, two prompts and one loop, and the diagram is what made the author see it that way.

The three phases are a requirements conversation, then two loop phases that share a mechanism. The requirements phase is interactive rather than automated: discuss the idea, identify the jobs to be done, break each job into topics, use subagents to pull information from links into context, and have a subagent write a spec file per topic. The output of phase one is a directory of spec files and nothing else. No code, no plan, no commits.

The loop then runs in one of two modes, and the mode is chosen by swapping which prompt file is present. In planning mode, the prompt's entire job is a gap analysis between the specs and the existing code, producing a prioritised task list and updating a single plan file. It is explicitly forbidden from implementing and from committing. In building mode, the prompt assumes the plan exists, picks a task, implements it, runs the tests, and commits.

Using the same mechanism for both is the design decision worth examining, and the guide gives four reasons. Building needs it regardless, because the work is inherently iterative and a fresh context per task is what provides isolation between them. Planning uses it for consistency, running through the same execution model even though it usually finishes in one or two iterations. The third reason is that a plan that turns out to need refinement can be revised by having the loop read its own output again. The fourth is simplicity: one mechanism for everything, clean file input and output, and the ability to stop and restart without special handling.

That last property is underrated. A two-mode design where the mode is a file swap is a system you can stop between any two iterations and understand, because the entire state of the run is on disk as plain text. That is what makes the rest of the design possible, and it is a better argument for the structure than the throughput numbers are.

## The ten-step cycle, and the parenthetical that keeps it from going in circles

The building mode is specified as a ten-step cycle, and reading it as a list is less useful than reading it as an answer to one question: how does a loop with no memory avoid redoing work?

The steps are: orient by having subagents study the spec files, read the plan, select the most important task, investigate the relevant source, implement with subagents doing the file operations, validate with a single subagent running the build and tests, update the plan marking the task done and noting anything discovered, update an agents file if there are operational learnings, commit, and then end the loop with the context cleared so the next iteration starts fresh.

Two steps carry the weight. Step four has a parenthetical instruction attached to it, and it is the most important sentence in the guide: when investigating, do not assume something is not already implemented. That is the entire defence against a stateless loop duplicating work. Each iteration arrives with a full window and no recollection of the last one, so the only way to know what exists is to read the source, and the parenthetical exists because that is precisely the step a fresh context is most inclined to skip in favour of writing the thing it was asked to write.

Step six is the other one. Validation is assigned to exactly one subagent, running the build and the tests, and it runs after implementation rather than alongside it. The guide calls this backpressure, which is the right word borrowed from the wrong domain: a single consumer downstream of many producers, whose result gates whether the iteration counts. With several subagents writing files and one subagent checking, the checking agent is the only part of the loop that can say no.

The final two update steps are the loop's only persistence beyond git. The plan file records what is done and what was discovered. The agents file records operational learnings, and it is updated conditionally rather than always, which is the difference between a notes file that grows without bound and one that stays worth reading. Between them, those two files plus the commit history are everything that survives a context clear, and that is a very small number of places for a system of this ambition. It is also exactly what makes the system debuggable: when a run does the wrong thing, the answer is in one of three text files.

## One sentence without the word and

The most immediately usable thing in the guide is a test for scoping a spec, and it takes one sentence to state.

The test is: can you describe the topic of concern in one sentence without joining two unrelated capabilities with a conjunction? A colour extraction feature passes, described as a system that analyses images to identify dominant colours. A user feature described as handling authentication, profiles and billing fails, and the guide annotates it as three topics. The stated conclusion is that if you need the word and to describe what something does, it is probably several things.

This is a good heuristic for a specific reason that the guide does not spell out. A spec file produces tasks, one spec produces many tasks, and the tasks are what the loop executes. A spec that quietly contains three capabilities produces three interleaved task streams, and a loop that can only hold one task in its smart zone will interleave them badly. The failure is not that the work is wrong, it is that each task is underspecified, because the spec was ambiguous about what done means. Catching that at spec-writing time, with a grammar test, is far cheaper than discovering it as a half-finished feature three iterations later.

The vocabulary around the test is what makes it usable rather than merely cute. There is a defined term for a high-level user need, a defined term for a distinct aspect within one of those needs, a defined term for the requirements document written per topic, and a defined term for the unit of work derived by comparing specs against code. The relationships are stated as a one-to-many chain: one need produces several topics, one topic produces exactly one spec, one spec produces several tasks, and the note that specs are larger than tasks is what prevents someone writing a spec per task and defeating the point.

The worked example is a mood board tool, decomposed into image collection, colour extraction, layout and sharing, each becoming a spec file and each becoming many tasks. It is a small example and it is enough, because the test does not care about the subject. If you cannot describe your feature without an and, you have a splitting problem, and no amount of prompt engineering fixes it.

## What the repository actually contains, and what it does not

There is no code here, and saying so plainly is the most useful thing this section can do.

The repository contains a readme, an HTML file that is the same document rendered as a formatted guide, a licence, a notice file, a gitignore, an editor configuration directory, a directory called files, and a directory called references. The references directory is where the provenance artefacts live: the diagram the method is built on, and the screenshot of the public disagreement that prompted the author to go back to the source material. That is a repository laid out like a piece of scholarship rather than like a package.

The files directory presumably holds the artifacts the method refers to, the prompt file and the agents file, since both are named as concrete paths in the guide and the directory name matches. Confirming that from the file listing alone would be guessing, and it does not matter much: the method is the deliverable, and the guide is specific enough to build your own files from.

What is absent is as informative as what is present. There is no tags, no releases, and the last commit to the default branch is dated at the start of March 2026. There is no contributing guide, no code of conduct, no security file, and no changelog. There is no issue template and no test of any kind, because there is nothing to test.

For a document like this, the absence of a release history is the right call. Pinning a guide to a version would imply that a later version supersedes an earlier one, and a guide that other people extend and argue with is better served by a commit hash and a date. A reader who adopts these practices should record the commit they read, because the arithmetic in the context section is the part most likely to be revised as models change, and it will be revised without a version marker.

The remaining question is what kind of adoption this invites. The guide is written as one person's reading of another person's method, with a link to the primary post that the originator gates behind a newsletter signup. Anyone who wants the original should go there rather than relying on a summary, which is the exact lesson the introduction is about.

## Conclusion

The Ralph Playbook is worth reading in full if you are running any kind of agent loop today, because the context arithmetic generalises to whatever model and harness you use and the one-sentence-without-and test is the cheapest quality gate in this genre. It is not a tool, so there is nothing to install and no code to audit, which means you should read it as an argument rather than a specification and check the numbers against your own harness rather than adopting them. If you try it, start with the planning mode and the topic scope test before writing a single prompt, because the failure mode of this whole approach is a loop that confidently re-implements what already exists, and the only defences it offers are the plan file, the agents file and reading the source before assuming.

## FAQ

### What is the Ralph Playbook?

It is a method guide for running autonomous coding loops, describing a three phase workflow that moves from an interactive requirements conversation through a planning loop to a building loop. The repository contains no code; it is documentation, published alongside a rendered HTML version of the same document.

### Why does the guide say context is the most important constraint?

Because it does the arithmetic: an advertised two hundred thousand token window is treated as roughly a hundred and seventy six thousand usable, with forty to sixty percent of that as the band where output quality holds. Running one tight task per loop iteration is recommended because it keeps the whole iteration inside that band, which makes task count per iteration a quality control rather than a throughput preference.

### What is the difference between the planning and building modes?

Both run the same loop and differ only in which prompt file is present. The planning prompt performs a gap analysis between the specs and the code and writes a prioritised task list, explicitly without implementing or committing. The building prompt assumes the plan exists, picks a task, implements it, runs the tests and commits.

### What is the one sentence without and test for?

It is a check on whether a spec covers one thing or several. If you need a conjunction to describe what a topic of concern does, it is probably multiple topics, and it should be split into separate spec files. The guide argues a test that passes is a spec that produces well-defined tasks, which is what keeps one task per iteration viable.

### How does the loop avoid reimplementing work it already did?

Because each iteration starts with a cleared context, the only defence is to read before writing. The guide instructs the investigating step not to assume something is unimplemented, and the loop's entire cross-iteration state is three things: the implementation plan file, an agents file for operational learnings, and the commit history.

## Sources

- [ClaytonFarr/ralph-playbook on GitHub](https://github.com/ClaytonFarr/ralph-playbook)
- [Issues](https://github.com/ClaytonFarr/ralph-playbook/issues)
- [License: MIT](https://github.com/ClaytonFarr/ralph-playbook/blob/main/LICENSE)
- [Project website](https://claytonfarr.github.io/ralph-playbook/)
- [README](https://github.com/ClaytonFarr/ralph-playbook/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/claytonfarr-ralph-playbook
