# dressage: agentic RL where compaction becomes training signal and six sandbox commands alias four functions

> A reinforcement learning framework for agents that call real tools, built on top of another project through a submodule and import-path hooks rather than a fork. Its two headline benchmark gains use different models and different metrics, and its optional storage component installs from a pinned git commit.

**Accio-Lab/Dressage** — Scalable RL for Any Agent and Sandbox.

- Repository: https://github.com/Accio-Lab/Dressage
- Stars: 787 · Forks: 93
- Language: Python
- License: Apache-2.0
- Published: 2026-09-18 · Updated: 2026-09-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/accio-lab-dressage

## The upstream framework is a submodule, and the extension points are import paths

The claim that this project is built on another rather than forked is specific and checkable. The upstream training framework sits in the repository as a git submodule, a submodule file is present at the root, and the integration happens through dotted import-path hooks rather than patched source. The listed integration points are rollout generation, reward processing, sample conversion and training behaviour, which is a large surface for a plugin mechanism, and the framing is that upstream receives no modifications. That trade is the interesting part of the design. A fork gives you freedom to change anything and costs you every upstream merge; hooks keep the tree clean and make you dependent on internals that upstream is free to move. The project also registers a plugin entry point for a separate orchestration package, so it is intended to be one component in a larger training system rather than a standalone tool, and its example directory includes a script demonstrating tool call hooks specifically.

## History compaction becomes training signal, and every segment becomes its own sample

The segment-aware training design is the most original idea here and worth reading twice. When a trajectory splits, because of history compaction or because the tool schema changed mid-run, the framework expands every segment into a separate training sample rather than discarding or truncating the leftovers. The terminal advantage from the anchor segment is then broadcast to its sibling segments, so each piece learns from how the trajectory actually ended. Segments are kept in the same training step by sharing identifiers for the rollout and the parent trajectory, and the denominators are normalised by equal prompts so that a trajectory which split five ways does not receive five times the gradient weight. What makes this unusual is the consequence: compaction, which in most agent frameworks is a context-management compromise forced on you, becomes a source of training data here, and a policy that triggers compaction early is generating extra signal rather than extra cost.

## Six sandbox command names are registered against four modules

The packaging registers one proxy command and then nine more around local sandboxes, and reading them closely shows duplicated wiring. There is a triple for a namespace-based local sandbox pool with start, status and stop, and then two more triples for blackbox clusters: one prefixed as local and one not. Of those two triples, the local start command and the local stop command both point at exactly the same functions as their non-local counterparts, so only the status commands differ. Six names therefore cover four modules, and two of them are exact aliases. Whether that is deliberate backwards compatibility for a renamed flag or an oversight is not stated. It matters in practice because a user who reads the documentation and runs the local-prefixed variant expecting different behaviour will get identical behaviour for two of the three operations, and the status output is the only place the difference would show up.

## The two headline results use different models and different metrics

Two recipes are announced with accuracy gains, and they cannot be compared to each other. The first trains a small four-billion-parameter model on a software-engineering dataset inside a Claude Code blackbox pipeline and reports accuracy on a verified software-engineering benchmark rising from 32.6 to 37.8 percent. The second trains a much larger mixture-of-experts model on a different dataset inside an OpenClaw blackbox pipeline and reports a strict pass metric on a separate evaluation rising from 41.4 to 55.4. Different base model, different dataset, different harness, different metric, and different definitions of what counts as a pass. Both are reported as point improvements with the arithmetic shown, which is honest presentation, and neither is a shared baseline. Two further reports are about infrastructure rather than model quality: a rebalancing change that reports higher effective throughput per GPU and shorter rollout time across different workload distributions, and a storage change that reports cutting peak trajectory memory on the master node by two thirds in a two-node experiment and by more than nine tenths in a 32-node capacity analysis.

## Staged grader injection is announced without saying when the grader becomes visible

One recipe is described as using staged grader injection, and the term is not explained anywhere in the visible text. In an evaluation-driven training loop, when a grader becomes visible to the policy is the difference between learning the task and learning the grader, so the staging schedule is not an implementation detail a reader can skip. The recipe is otherwise described carefully, with fresh-sandbox evaluation and anti-cheating safeguards, so the safeguards are clearly part of the design; what is missing is the mapping between the two, namely what the stages are and what the policy can observe at each one. The other recipe in the same family is described as a reproducible 64K setup, which tells you the context length and nothing about the grading schedule. Treat both numbers as results of a specific harness configuration rather than as properties of the framework.

## Two dependencies are pinned exactly and one extra installs from a git commit

The dependency list mixes ranges with exact pins, and the pins are the interesting part. Two of the seven runtime dependencies are fixed to a single version, the distributed computing library and the model library, while the web framework, the HTTP client, the validation library, the ASGI server and the remote sandbox client all take ranges. Pinning the two that dominate memory and numerics is a defensible call for a training framework. The optional storage extra is a different kind of decision: it installs a component directly from a git repository at a fixed commit rather than from a published artefact with a hash. That is common in research software and it means the dependency is only as stable as that commit's hosting, and that a vulnerability in it arrives without a release cycle. Review the commit before you add it to anything with credentials.

## The default test run excludes the tests that cost money

The test configuration is worth one look because it tells you what the maintainers consider expensive. The suite runs from a single test directory, and the default options exclude one marker entirely, described as covering end-to-end tests that are paid or environment-dependent across the orchestration package, GPU hardware and the remote sandbox service. So by default you get the tests that run anywhere, and the ones that need a rented sandbox or a specific accelerator are opt-in. That is the right default and the marker name says so plainly, which is more transparency than most repositories manage. The project also separates its own concerns physically: a top-level directory for the blackbox server, which is also published as its own package and pinned as an extra, a docker directory, a docs directory with bilingual reports, and an examples tree containing data, orchestration examples, shell scripts and a single hook demonstration.

## Conclusion

Evaluate this as research infrastructure with a short public history, not as a stable library. The project announced itself open source in June 2026, declares version 0.1.0, has published no releases, and its last commit is dated 2026-08-27, so the interface you read may not be the interface next month. What it does have is a clear architectural bet worth understanding: agent semantics separated from execution placement, so a Python tool loop, an HTTP agent, a local namespace sandbox and a remote one all converge on the same path to training. Three things to weigh. Its extension model depends on another project's internals through import-path hooks and a submodule pin, which is cheaper than a fork and more fragile to upstream change. Its storage extra installs a dependency straight from a git commit rather than a published artefact, so review that pin before using it. And the two published results measure different models against different metrics, so neither is a baseline you can compare your own run against.

## FAQ

### What is the Dressage framework?

An agentic reinforcement learning training framework built on top of slime, bridging policy rollouts, sandboxed tool execution and training data conversion through a shared proxy layer. It trains agents that call real tools, such as code editors, shell commands, file access and retrieval interfaces, in both whitebox Python loops and blackbox HTTP agents through one interface.

### Does Dressage fork the slime framework it builds on?

No. Slime sits in the repository as a git submodule, and rollout generation, reward processing, sample conversion and training behaviour plug into upstream through dotted import-path hooks rather than patched source, so no fork is maintained.

### How does Dressage handle token drift between turns?

It records training evidence at token granularity, including token id, logprob, loss mask, token version and token expert, and encodes only each turn's append delta, splicing token ids incrementally, which the project says avoids retokenization drift. The same record drives pause and resume at token boundaries and routing replay.

### What results does Dressage report?

A small model trained on a software-engineering dataset improved accuracy on a verified benchmark from 32.6 percent to 37.8, and a larger mixture-of-experts model improved a strict pass metric from 41.4 to 55.4 on a different evaluation. Infrastructure changes report higher throughput per GPU and lower peak trajectory memory.

### How are Dressage's optional dependencies installed?

Through three extras: a test extra with the test framework, a storage extra that installs a component directly from a git repository at a fixed commit, and an orchestration extra pinning an orchestration package and a blackbox server package. Two runtime dependencies are pinned to exact versions while the rest take ranges.

## Sources

- [Accio-Lab/Dressage on GitHub](https://github.com/Accio-Lab/Dressage)
- [Issues](https://github.com/Accio-Lab/Dressage/issues)
- [License: Apache-2.0](https://github.com/Accio-Lab/Dressage/blob/main/LICENSE)
- [README](https://github.com/Accio-Lab/Dressage/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/accio-lab-dressage
