# inspect-robots: pinning a version matters more here than usual

> inspect-robots is a MIT-licensed evaluation framework for physical AI that runs an LLM agent or a VLA policy against a real arm, a humanoid or a simulator, with grader scores, transcripts and full configuration in every log. It is also a 0.x project that tells you to pin a version before depending on it.

**robocurve/inspect-robots** — Open source evals for physical AI. Run any LLM/VLA on any arm/humanoid against any real/sim benchmark.

- Repository: https://github.com/robocurve/inspect-robots
- Website: https://inspectrobots.org
- Stars: 644 · Forks: 77
- Language: Python
- License: MIT
- Published: 2026-09-13 · Updated: 2026-09-13 · Language: en
- Canonical page: https://hysenlabs.com/projects/robocurve-inspect-robots

## The API may change between releases, so pin a version

The first note in the project documentation is the one to read before anything else: this is early development and the API may change between releases, so pin a version before depending on it. The version numbers explain why. The latest three releases are v0.60.0, v0.59.0 and v0.58.0, published on 2026-09-30, 2026-09-22 and 2026-09-02, which is a cadence fine for a library that owns its own breaking changes but hostile to an unpinned dependency in someone else's robot stack. The metadata is consistent with the warning rather than contradicting it. The classifier says Development Status :: 3 - Alpha, the audience is Science/Research, and the version is produced dynamically by hatch-vcs from tags rather than written into pyproject.toml, which is another reason a loose install is a bad idea. Supported interpreters start at Python 3.10 and the classifiers run through 3.13. The license is MIT with the license file declared, and there is a CITATION.cff at the root, so results are meant to be referenced rather than quoted from a chat window.

## The core is numpy plus the standard library, everything else is an extra

The dependency design is unusually disciplined for a project at version 0.60. The core installs with:```bash
uv venv && uv pip install "inspect-robots[rerun]"
```The `rerun` extra exists to power the live run viewer, and the project states plainly that the core itself is numpy-only. That constraint shows up in the optional dependency table. `rerun` pulls the rerun SDK and pillow, `viz` declares exactly the same pair, `docs` pulls griffe, `dev` adds pytest, coverage, ruff, mypy and pre-commit, and `all` resolves to the `rerun` extra. Having `viz` duplicate `rerun` byte for byte is a small redundancy, and it is the kind of thing that either gets cleaned up or gets quietly depended on. The dev extra contains the most interesting line in the file, pinning `numpy<2.5` with a comment explaining that numpy 2.5 type stubs use 3.12-only syntax that mypy rejects when the check runs against Python 3.10, and that the pin exists for a deterministic dev and CI gate while the runtime stays free. Someone will hit that again on a fresh environment.

## uv run inspect-robots silently uninstalls the plugin you just added

One operational warning is specific enough to be worth repeating verbatim in spirit, because the failure it prevents looks like a bug in the plugin rather than in your workflow. Inside an existing uv project, the documentation tells you to avoid `uv run inspect-robots`, since that form re-syncs the environment to the lockfile and silently removes whatever `uv pip install` just added. Since every capability in this project arrives as a plugin, the effect is that the CLI returns without the policy you just installed and gives no obvious reason why. The recommended alternative is dull and correct: any venv workflow works, activate the environment once with `source .venv/bin/activate`, or the Windows equivalent `.venv\Scripts\activate`, and then call `inspect-robots` directly so the executable resolves inside the environment you actually built. The repository ships a `uv.lock`, so teams arriving from Python will reach for `uv run` by habit. The plugin list is long enough that this bites more than once.

## The setup wizard asks only what the embodiment plugin declares

Configuration is generated from the embodiment rather than from a form the framework invented, which is the detail that makes adding a new arm tractable. Install the plugin for your rig and run setup once:```bash
source .venv/bin/activate
uv pip install inspect-robots-yam   # provides the molmoact2 policy + yam_arms rig
inspect-robots setup
```The `inspect-robots-yam` plugin is the worked example, bringing the `molmoact2` policy and the `yam_arms` rig with it. What `setup` does is bounded: it picks your defaults, finds your cameras, and asks about the behavior toggles and numeric settings the embodiment plugin declares, with `auto_start` on yam given as the example. It then writes `~/.config/inspect-robots/config.ini`. On a different rig the same wizard runs after installing that rig's plugin, and you type the component names at the prompts instead. If you would rather not use the wizard, the CLI guide covers writing the config file by hand, which matters for fleets where the file is generated per machine rather than per user. The pattern is worth noting because it keeps the framework from needing to know anything about your hardware.

## molmoact2 is only a client, and the server does not survive a reboot

Several policies in this ecosystem are clients to something that has to be running elsewhere, and the documentation is direct about the failure mode. Nothing moves until the MolmoAct2 server is listening, and that server does not start itself and does not survive a reboot. The bring-up runs on the GPU machine from the MolmoAct2 repository:```bash
# On the GPU machine, from the MolmoAct2 repo. Leave it running, e.g. in tmux:
python examples/yam/host_server_yam.py --host 0.0.0.0 --port 8202
curl http://127.0.0.1:8202/act      # 200 means the server is ready
```The curl probe matters more than it looks, because a listening port and a ready endpoint are different states and the project tells you which response means ready. The comment about tmux is not decoration either; a foreground server on a robot that reboots is a session you will restart by hand. The alternative is a different class of policy. In-process policies, `agent` being one and the mock `scripted` another, need no server at all, which is what makes them the sane default for a first run and for CI. On a rig other than yam you start whatever serves your policy, and the policy slot stays the same either way, which is the property that makes a benchmark portable across embodiments.

## One approver-checked motion chunk per tool call

The agent policy is where the safety story is most concrete, and it is not a general-purpose agent with robot access. With the agent plugin installed, a frontier LLM drives the same rig through tool calls, and the documented granularity is one approver-checked motion chunk per call. The run itself looks like this:```bash
inspect-robots "place the fork on the plate" --policy agent \
    -P model=anthropic/claude-fable-5 -P effort=low
```The `-P` flags are the policy parameters, so `model` and `effort` are passed the same way as any other plugin setting. Credentials come from a `.env` file in the working directory, which the CLI loads by itself, with real environment variables taking precedence over the file, and `.env.example` is the template. What a run leaves behind is the reason the project exists: grader scores, the LLM transcript and the full configuration, plus a replayable `.rrd` file beside the eval log holding the cameras, proprioception and actions streamed from the pipeline. `inspect-robots inspect LOG.json --transcript` reads the conversation back, `inspect-robots view LOG.json` opens the HTML report, which for a `--store-frames` run includes the frames the model actually saw, and `inspect-robots view logs/` builds a browsable index once several runs accumulate.

## prior_learnings is read once and pinned by content hash

Retrying a failed run with what the failure taught you is a two-command loop, and the way the second run refers to the first is the interesting part:```bash
inspect-robots summarize logs/failed-run.json
inspect-robots "place the fork on the plate" --policy agent \
    -P prior_learnings=logs/learnings/failed-run.md
````summarize` writes the markdown file, and `prior_learnings` points the next run at it. The policy reads that file once, and the resolved path together with the content hash is recorded in the eval configuration. That detail is what keeps the loop honest. An eval whose configuration says which notes were used, by path and by hash, can be compared against a later run whose notes differ, because you can tell whether an improvement came from the model, from the embodiment, or from someone editing a summary file between runs. Without it, an agent benchmark is a moving target. The same care shows up in the viewing commands, where the report rather than the terminal is the artifact of record, and in the CLI overrides such as `--no-rerun-save`, `--no-rerun`, `--no-store-frames` and `--max-steps 300` that let a recorded run be repeated without paying for the same storage twice.

## One plugin says pip install where every other says uv pip install

A small inconsistency in a README this detailed is worth naming, because it is the kind of thing that costs a reader ten minutes. Every plugin section installs with `uv pip install`, and the voice section alone says `pip install`:```bash
pip install inspect-robots-voice
inspect-robots run --policy agent -P model=anthropic/claude-opus-5 \
    --voice \
    --instruction "place the fork on the plate"
```Beyond the install line, the voice plugin has the strictest privacy claim in the project. It keeps the microphone open for the whole run and delivers each spoken remark to the policy at its next inference, transcribed locally with no keys and no network. Silence sends nothing at all. The limits are explicit too: voice is feedback only, so ending an episode and recording verdicts stay on the keyboard, and typed console feedback keeps working alongside it with both landing in the transcript and the eval log with their source recorded. The other caution in the same area comes from the code-as-policy plugin, which evaluates Python the model wrote against SAM3 segmentation, Contact-GraspNet planning, Pyroki inverse kinematics and speed-limited joint-motion helpers. That plugin README is where the model-code trust boundary is documented, and it is the page to read before pointing generated code at real hardware.

## Conclusion

inspect-robots is worth adopting if your team already runs robot policies and has no way to compare them, because the combination of a fixed embodiment, a fixed instruction and a log that holds the transcript, the grader scores and the whole configuration is the thing that makes a comparison defensible. Four cautions. It is 0.x with a stated API that can change between releases, so pin the version and re-read the changelog on upgrade. Inside a uv project, `uv run inspect-robots` re-syncs to the lockfile and removes your plugin, so activate the environment and call the binary. Several policies are clients to a server that does not start itself or survive a reboot. And the code-as-policy plugin evaluates Python the model wrote, which is a trust boundary you should read about in that plugin before pointing it at real hardware.

## FAQ

### How do I install inspect-robots with the live run viewer?

With uv venv and uv pip install for the rerun extra, which pulls the rerun SDK and pillow. The numpy-only core installs the same way without the extra, and inside a uv project you should activate the environment and call inspect-robots directly rather than using uv run.

### Why does nothing move when I run inspect-robots with the molmoact2 policy?

Because molmoact2 is only a client. The MolmoAct2 server has to be listening on the GPU machine, it does not start itself, and it does not survive a reboot, so you start host_server_yam.py and confirm the act endpoint returns 200 before running the eval.

### What does an inspect-robots log contain?

Grader scores, the LLM transcript and the full configuration, plus a replayable .rrd stream of cameras, proprioception and actions saved beside the eval log, which you can read back with inspect and view as an HTML report.

### Can an inspect-robots run learn from a previous failure?

Yes. inspect-robots summarize writes a markdown file from a failed log, and the next run takes it through prior_learnings, which the policy reads once while the resolved path and content hash are recorded in the eval configuration.

### Is inspect-robots safe to depend on without pinning a version?

The project states it is in early development and that the API may change between releases, so pin a version before depending on it. The latest releases are v0.60.0, v0.59.0 and v0.58.0, and the package is classified as alpha.

## Sources

- [License: MIT](https://github.com/robocurve/inspect-robots/blob/main/LICENSE)
- [Project website](https://inspectrobots.org)
- [README](https://github.com/robocurve/inspect-robots/blob/main/README.md)
- [Releases](https://github.com/robocurve/inspect-robots/releases)
- [robocurve/inspect-robots on GitHub](https://github.com/robocurve/inspect-robots)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/robocurve-inspect-robots
