# statewright/statewright: the benchmark headline claims two conditions the table only records one

> A deterministic Rust state machine that restricts which coding tools an agent may call in each phase, enforced across several agent hosts with per-state model routing. The research section leads with a before-and-after result on a five-task slice of a 2,294-instance benchmark, while the table beside it has a single pass column and no baseline, and the quickstart requires a hosted account key even though the repository ships a self-hosted stack.

**statewright/statewright** — State machine guardrails for AI agents

- Repository: https://github.com/statewright/statewright
- Website: https://statewright.ai
- Stars: 504 · Forks: 23
- Language: Rust
- License: NOASSERTION
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/statewright-statewright

## The headline is a before-and-after the table never shows

The research section opens by saying two local models went from two of ten attempts passing to ten of ten with the constraints applied, on the same tasks and same hardware. The table below it does not contain that comparison. It has one column for a 26-line bug fix and one column for a five-task slice, each holding a single pass or fail, with no baseline condition recorded anywhere. So the strongest claim in the file cannot be checked against the evidence printed next to it. The caveats are at least disclosed rather than buried. The benchmark is named as a five-task subset and explicitly not the full 2,294-instance suite. One passing result is footnoted as requiring a specialised line-editing tool adaptation. One row is marked as tested on two of the five tasks and footnoted as added after the initial run. And the largest model is credited with two of two, fewer tasks than the models above it.

## A thirteen gigabyte floor is attributed to the model, not the tool

The claim is that guardrails help at any model size, but with a stated threshold. Below about 13GB, models can still emit tool calls yet cannot hold enough file content in context to make accurate edits, and above that threshold the guardrails are described as starting to turn failures into completions. The failure mode is diagnosed specifically: smaller models identify the bug correctly but cannot serialise a surgical edit and rewrite the entire file instead. The attribution is explicit, saying this is a model limitation rather than a limitation of the guardrails. What the threshold means in practice is that the tool's value is concentrated above a hardware floor, so the 3.3GB and 7.2GB rows in the table fail regardless of how the workflow is configured. A team running local models below that line is being sold a mechanism whose benefit the project's own data says it cannot deliver.

## Three different host counts appear in three places

The introduction says a workflow is enforced across five named hosts. The quickstart then offers four install paths, three of which fetch a distinct scoped package through a package runner, for example:

```bash
npx statewright-codex@latest init
```

The fourth installs through a plugin marketplace with an add step followed by an install step. There is no entry for one of the five hosts. The architecture section then names six hosts, adding one that appears in neither of the earlier lists. The mismatch is easy to miss because every individual path is correct, and it matters most for the host present in the architecture section but absent from the quickstart, since a reader has no documented way to install it. Any comparison of host support should be made against the architecture list, which appears to be the complete one.

## The quickstart needs a hosted key while the repository ships self-hosting

After the install, the documented next step is that your browser opens, you sign up on the project's hosted site, generate a key and paste it. That is the path presented as the way to finish setting up, and it applies to the package-runner installs. Meanwhile the repository contains a self-hosted directory, a compose file, and a section of the architecture describing an executor that owns credentials, delivery and telemetry. So there are two deployment models, and the one you are walked through first requires an account on someone else's service. The compose file makes the alternative concrete: a message broker with its own persistence mode enabled and a relational database, both with credentials set inline to a fixed development value and both with their ports published to the host, including the broker's monitoring port. Nothing in the visible documentation explains which features differ between the two models.

## Three model identifier formats share one workflow example

The routing example assigns a different model to each of three states and inherits a default for a fourth case. The three names do not share a format. One is a bare provider-neutral name with a hyphenated version and a date appended. The second inserts a dot between version components and also appends a date. The third carries an explicit provider prefix and a dotted version with no date at all. A consumer writing a workflow has to guess which convention their host expects, and the documentation does not say. There is a second configuration surface as well: the workflow file selects a model per state through a model field, while the agent binary separately accepts a config file with a routing block for per-state endpoint, temperature and context window overrides. Which one wins, and whether they compose, is not stated.

## Model routing means something different on every host

The routing section is candid about this in a way few tools are. The executor is said to use the strongest deterministic boundary each host exposes, and then it enumerates what that means per host. Two hosts switch models live. Two others resume the same session after a route change, which is not the same as switching and means conversation state carries across the boundary. One applies the route at startup only, so a workflow that escalates mid-run cannot change models on that host. One more routes on each turn through a dedicated adapter. So the same workflow file produces four different runtime behaviours depending on where it runs. The per-state model idea is the strongest feature described in the file, and it is also the one whose guarantees you cannot assume transfer between hosts.

## No runtime dependencies for the engine, two services for the deployment

The engine is described as a pure Rust state machine evaluator with states, transitions, guards and tool restrictions, characterised as deterministic, with no model in the loop and no runtime dependencies. That is a genuinely strong property for something meant to sit in front of an agent, because the part that enforces policy cannot be influenced by the thing it constrains. The deployment story is heavier. A self-hosted install brings up a message broker with persistence enabled and a database, and the compose file wires both with inline credentials, a named volume for each, and host port mappings for the broker's client port, the broker's monitoring port and the database's port on a non-default host number. Those are development defaults in a file whose name does not say development, which is the kind of thing to check before pointing at anything shared.

## Two version lines, one set of tags, and an unnamed licence

The workspace manifest declares version 0.1.0 and an Apache-2.0 licence, while the three releases in the repository are all prefixed for one specific host plugin and are numbered 0.3.1 through 0.3.3. So the tags track the plugin, the core workspace sits three minor lines behind at 0.1.0, and there is no tag a user of the engine can pin to. The licence is a second loose end: no licence value appears in the repository metadata at all, while the manifest declares one and both a licence file and a separate patents document sit at the root. The workspace also lists five member crates, one of which does not appear in the three-layer architecture description, so the layering in the documentation is not the whole build. None of this is disqualifying, but all of it is the kind of detail to confirm before depending on the project.

## Conclusion

This suits someone whose agent keeps calling the wrong tool at the wrong moment and who wants a hard stop rather than a better prompt. Evaluate the research claims on their own terms, since the headline is a before-and-after figure the table does not show and several rows carry footnotes about adjusted tooling. Verify the model routing semantics for your specific host, because a live switch and a session restart are not the same behaviour, and decide early whether you need the hosted key or the self-hosted stack.

## FAQ

### What does Statewright actually restrict?

Which tools an agent can see and call in each phase of a workflow. A planning state gets read-only tools, implementation unlocks edit tools with limited shell access, and write-via-redirect and destructive operations stay blocked even when shell access is allowed. Testing permits only designated commands. Calling a tool outside the current phase is rejected with a message naming what is available and how to transition.

### How do I install Statewright?

Three hosts install through scoped packages run with a package runner, each with its own init command, while one installs through a plugin marketplace with an add step followed by an install step. After installing, the documented step is that a browser opens so you can sign up on the hosted site, generate a key and paste it. A self-hosted stack is shipped as an alternative.

### Which model sizes benefit from Statewright's guardrails?

The project states the floor is around 13GB, below which models can emit tool calls but cannot retain enough file content for accurate edits and rewrite entire files instead. The 3.3GB and 7.2GB rows in the research table fail, and the passing rows begin at 13.8GB. The limitation is attributed to the model rather than the guardrails.

### Does per-state model routing work the same on every host?

No, and the documentation says so. Two hosts switch models live, two resume the same session after a route change, one applies the route at startup only, and one routes on each turn through a dedicated adapter. The workflow file selects a model per state, and the agent binary separately accepts a config file with per-state routing overrides.

### What evidence does Statewright give for its research claims?

A five-task subset of a 2,294-instance benchmark, presented in a table with one pass column and no baseline, while the surrounding text claims a rise from two of ten attempts passing to ten of ten. Footnotes disclose that one pass needed a specialised line-editing tool adaptation and that one row was added after the initial run. A research brief is linked from the project's site.

## Sources

- [Issues](https://github.com/statewright/statewright/issues)
- [Project website](https://statewright.ai)
- [README](https://github.com/statewright/statewright/blob/main/README.md)
- [Releases](https://github.com/statewright/statewright/releases)
- [statewright/statewright on GitHub](https://github.com/statewright/statewright)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/statewright-statewright
