# StateM's default edge retry is unbounded, and its only release is a benchmark bundle

> henryqin1997/statem is a dependency-free Python CLI that stores an agent's procedural state outside the model context, as a versioned graph of states, transitions, and executable checks. Its edge schema treats retry as unbounded unless a ceiling is set, its example runbook shows one field in two different shapes, and its 88.8 percent figure is descriptive accuracy on named benchmark subsets.

**henryqin1997/statem** — CLI runbook for agent long run.

- Repository: https://github.com/henryqin1997/statem
- Stars: 1,286 · Forks: 116
- Language: Python
- License: Apache-2.0
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/henryqin1997-statem

## An edge without max_attempts retries without a ceiling

The edge schema has a field called max_attempts, described as an optional positive retry ceiling for that edge and source-node entry.

And then the sentence that matters: leaving it out preserves the default unbounded retry behavior.

So a runbook that describes a repair loop, which the overview diagram does, and that does not set max_attempts on the repair edge, will retry indefinitely. The project states this rather than hiding it, which is the right thing to do, and the mechanism is visible in the transition description as well: a blocking check leaves the agent in the current state with the failure recorded for repair. That is correct behaviour for a repair loop. Combined with an unbounded edge, it means nothing in the schema forces the loop to close.

What the field does when set is narrower than the name suggests. Each real goto consumes one attempt, blocked checks count, and the attempts are scoped to that edge and that source-node entry rather than to the node or the run. So two different edges leaving the same node get separate budgets, and re-entering the node from elsewhere starts again.

For anyone writing a runbook, max_attempts is the field that decides whether a stuck agent keeps trying or stops and reports.

## The example runbook shows one gate field in two different shapes

The minimal runbook is the clearest document in the repository, and it has a formatting inconsistency in the middle of it.

The plan node writes its before_transfer as a mapping, with a type of checklist and an items list beneath it. The execute node writes its before_transfer as a sequence, with a dash for a command entry followed by a dash for a checklist entry, and that command runs the test suite.

Both are gate declarations and both are in the same document, separated by one node. So either the field accepts a single gate or a list of gates, in which case the mapping form is a one-element list written the other way, or there are two spellings and the reference does not say which is which.

It is the kind of ambiguity that a validator would normally settle, and there is a validate command for exactly that purpose. The included example runbook passes it, since it is shipped as a working example, so whatever shape it uses is at least accepted.

The example also stops partway. The handoff node's prompt reads Summarize th and ends there, so the last node in the minimal runbook has no complete instruction.

The reference section on max_attempts is cut off in the same way, ending after the words and a fresh.

## Four layers, and only one of them belongs in git

StateM's central design decision is a split between what is shared and what is per run, and the table that expresses it has four rows.

The static runbook holds nodes, edges, prompts, hooks, and gates, and it is the layer you commit. The runtime state holds the current node, the history, results, and timestamps, and it is not committed. Dynamic checks are task-specific verification for the current entry, also not committed. Durable project notes, meaning plans, decisions, progress, and artifacts, are usually committed.

So a repository carries the procedure and the durable thinking, while the machine carries the position and the record. The point of the runtime state living outside git is that a run can survive a disposable checkout, which is what the STATEM_STATE_DIR setting is for, since runtime data defaults to a .statem/ directory inside the working tree.

The dynamic checks layer is the interesting one and it has a constraint built into its description: an agent can register task-specific checks for the current state entry without mutating the shared runbook. That is what keeps the committed procedure stable while still letting one run verify something specific.

The package itself ships only the core. The packaging configuration includes packages matching statem and excludes plugins, integrations, examples, and tests, so the distribution is the CLI rather than the runbooks people write.

## A transition is six steps and a failed gate leaves you in place

The transition is described as a transaction, which is the right framing, and it has six steps in order.

First the requested outgoing edge is resolved. Then the current node's before_transfer checks run. Then the current-entry dynamic checks are loaded and run. Then the edge's condition is evaluated. Then the current node's out_hook and the edge hook run. Finally the transition is recorded, a new target entry is created, and the target node's in_hook runs.

The ordering is the interesting part. Both the shared gates and the run-specific gates are evaluated before anything is written, and the hooks that persist progress run after those gates have passed but before the transition is recorded. The edge hook is described as prepare-transfer work, and the target in_hook as the place to load context or initialise evidence.

What happens on failure is stated plainly: if a blocking check fails, the agent remains in the current state with the failure recorded for repair.

So the failure path is not an exception that unwinds the run. It is a recorded, resumable fact that the next attempt can act on, which is the whole reason the state is kept outside the context in the first place.

## The benchmark runbooks are named subset and family, and 445 is not the whole suite

The single release in this repository is tagged deepseek-policy9-tb21-artifacts-20260818 and titled as a DeepSeek V4 Flash policy-v9 set of Terminal-Bench 2.1 artifacts. There is no version tag, and the project metadata says 0.1.0.

The runbooks that release refers to are in the examples directory, and their filenames are precise about scope. One is a server readiness subset, as a YAML runbook and a Markdown companion. The other is a git webserver deploy family, again as a pair. The word subset appears in one name and family in the other, and neither is named as the full benchmark.

That matters when reading the headline number. The reported figure is 88.8 percent descriptive accuracy on Terminal-Bench 2.1, given as 395 out of 445 trials, with 88.76 percent unrounded. So 445 trials is the denominator, across two named families of tasks rather than the entire suite, for one model under one policy revision.

Descriptive accuracy is also not the same word as pass rate. Nothing in the visible text defines it further, and the exact trial count is given alongside it, which is the more checkable half of the claim.

The same date carries the other news item: the paper ranked first on the Hugging Face daily papers listing on 2026-08-18, and a project site, a paper PDF, and a forty-second demo are linked from the header.

## 88.8 percent is descriptive accuracy for one model on one policy

It is worth separating what the number is from what it is not, because the phrasing in the announcement is careful and the shorthand is not.

The figure is descriptive accuracy, not a task completion rate, on Terminal-Bench 2.1, for the DeepSeek V4 Flash model, under policy revision nine, over 395 successful trials out of 445, with the unrounded value given as 88.76 percent so the rounding can be checked.

Two things are pinned down tightly. The denominator is exact, and the rounding is disclosed. That is a better practice than a bare percentage.

Two things are not pinned down at all in the visible text. What descriptive accuracy measures relative to task success, and whether the 445 trials are the complete set those two runbook families define or a selection from them.

The filenames support the narrower reading, since one is called a subset and one a family. A subset inside a benchmark of that name means something was left out, and the release artifacts sitting in examples/ rather than in a results directory suggests the runbooks themselves are the artifact being shared.

So the honest way to repeat the claim is: one model, one policy, two named task families, 445 trials, 88.76 percent by a metric the repository defines for itself.

## The package declares no dependencies while the reference offers an LLM review gate

The zero runtime dependencies claim is not marketing, it is in the packaging metadata. The dependencies field is an empty list, and the core requires Python 3.11 or newer.

The install is a clone and an editable install, with no package index step and no lock file involved:

```bash
git clone https://github.com/henryqin1997/statem.git
cd statem
python3 -m pip install -e .
```

That is affordable precisely because nothing has to be resolved.
That is a real constraint on what the tool can do, and one part of the feature list runs into it. The executable transition gates are described as checklists, commands, predicates, manual approval, and LLM review. The first four need nothing beyond a shell. The fifth needs a model.

With an empty dependency list there is no model client in the package, so an LLM review gate has to be arranged some other way: a command in the gate list that calls out to something the operator provides, or the agent running the review itself as part of its own turn. Neither is stated in the visible text.

That is not a defect. It is the usual trade for a tool that installs inside somebody else's agent environment, where the model is already there. It does mean the one gate type with an external dependency is also the one whose mechanism is least visible.

The rest of the packaging is plain. A single console entry point pointing at the CLI main function, a package find that includes only the core namespace, and a setuptools build requirement. The classifier list names the license and Python 3 and 3.11, with no entries for later 3.x releases even though the floor is 3.11 or newer.

## The comparison table gives a TODO list the same marks twice

The StateM case is built on a table comparing five approaches across five properties: whether the approach remembers the phase, blocks invalid transitions, supports repair loops, survives a context refresh, and is agent-editable.

Read the rows rather than the last one. A prompt-only workflow gets partial, no, informal, no, yes. A TODO list gets partial, no, informal, yes, yes. A CI pipeline gets yes, yes, limited, yes, and usually no. A general workflow engine gets yes, yes, yes, yes, and rarely.

StateM's row is five yeses.

Two observations follow. A TODO list matches StateM exactly on two of the five columns, surviving a context refresh and being agent-editable, so the differentiator is not the artifact an agent carries but what enforces the transitions. And a CI pipeline matches on four of five, with only agent-editable separating them, which is a fair statement about CI rather than an argument against it.

The table is self-authored, and its honest weakness is that the yes and partial cells are not defined anywhere. Nothing says what fraction of phases a partial score means.

The project also positions itself by size rather than by capability, stating that it is deliberately smaller than a workflow engine and is a state-aware runbook an agent can read, author, inspect, and repair from the command line. The design puts that same shape on screen, with four states and a repair loop back to execute.

## Conclusion

StateM fits someone whose agent runs long enough that progress keeps ending up only in chat history, and who wants the phase to survive a context refresh. Four things to check before you adopt it. Whether your loops terminate, since retry is unbounded unless you set a ceiling on the edge, and a repair loop without one can run indefinitely. That your YAML matches your expectation of the schema, since the reference shows one gate field in both a mapping form and a list form. What the 88.8 percent covers, which is descriptive accuracy for one model on one policy over two named subsets rather than a pass rate on the whole suite. And what an LLM review gate calls, given the package declares no dependencies at all.

## FAQ

### What is StateM and what problem does it solve?

It is a command-line state machine for long-running AI agents. It moves procedural state out of the model context into a versioned runbook, so planning, execution, verification, repair, and handoff stay separate instead of collapsing into one long prompt, and a new session can reconstruct what happened.

### Does StateM have any runtime dependencies?

The dependency list in the project metadata is empty, and the core package requires only Python 3.11 or newer. It is installed in editable mode from a clone and exposes a single console entry point. The feature list does include an LLM review gate, whose mechanism is not declared in the metadata.

### How do I stop a StateM repair loop from retrying forever?

Set max_attempts on the edge. The field is an optional positive retry ceiling for that edge and source-node entry, and leaving it out preserves the default unbounded retry behavior. Each real goto consumes one attempt and blocked checks count.

### What does StateM record when a check fails?

The agent remains in the current state with the failure recorded for repair. A transition resolves the edge, runs the node gates, runs the dynamic checks, evaluates the edge condition, runs the out and edge hooks, and only then records the transition and runs the target in_hook.

### What is the StateM Terminal-Bench 2.1 result?

The release reports 88.8 percent descriptive accuracy, given as 395 out of 445 trials with 88.76 percent unrounded, for DeepSeek V4 Flash under policy revision nine. The two runbooks in the repository are named a server readiness subset and a git webserver deploy family.

### Which StateM files go into version control?

The static runbook, meaning nodes, edges, prompts, hooks, and gates, is committed. Runtime state, meaning the current node, history, results, and timestamps, is not, nor are dynamic checks. Durable project notes such as plans, decisions, progress, and artifacts are usually committed.

## Sources

- [henryqin1997/statem on GitHub](https://github.com/henryqin1997/statem)
- [Issues](https://github.com/henryqin1997/statem/issues)
- [License: Apache-2.0](https://github.com/henryqin1997/statem/blob/main/LICENSE)
- [README](https://github.com/henryqin1997/statem/blob/main/README.md)
- [Releases](https://github.com/henryqin1997/statem/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/henryqin1997-statem
