# ClawProBench grades agents in a live runtime, and admits the open dataset can be gamed

> A benchmark harness with 102 active scenarios out of a 164-entry catalog, deterministic grading and two vendored harnesses for cross-evaluation. Its own changelog says a fully open benchmark cannot avoid vendors optimising for it, which is why there is a second, closed dataset and a separate report-sourced scoreboard.

**suyoumo/ClawProBench** — ClawProBench is a live-first benchmark harness for evaluating LLM agents   in the OpenClaw runtime with deterministic grading and repeated-trial   reliability.

- Repository: https://github.com/suyoumo/ClawProBench
- Website: https://suyoumo.github.io/bench/
- Stars: 824 · Forks: 54
- Language: Rust
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/suyoumo-clawprobench

## 102 active scenarios out of a 164 entry catalog

The size of the suite is reported by the harness itself rather than asserted in prose, and the two commands that report it are the quickest way to see what is in a checkout:

```bash
python3 run.py inventory --json
python3 run.py inventory --benchmark-status all --json
```

The first gives the worktree inventory, the second widens it. The numbers are 102 active scenarios, 164 total catalog scenarios, and 62 of those incubating, which means the difference between what a leaderboard is computed from and what exists as material. Coverage is then sliced by profile. `core` is the default ranking path, and `intelligence`, `coverage`, `native` and `full` trade a longer run for broader active coverage. That structure is the first thing to understand about the project: a benchmark where the active set is a subset of the catalog, where 62 scenarios are still incubating, and where the number in the leaderboard depends on which profile produced it.

## A Rust project whose Python surface is three dependencies

The repository is recorded as Rust, and the Python side is deliberately thin. The whole requirements file is:

```
PyYAML>=6.0
fastapi>=0.110
uvicorn>=0.29
```

Three packages: YAML parsing, an ASGI framework and its server. That tells you what the Python entry point does. It reads configuration and serves something over HTTP, and the inventory command runs through `run.py` at the root. Everything else is directories with one job each: `scenarios/` for the tasks, `fixtures/` for their inputs, `custom_checks/` for the graders, `mock_tools/` for the tool surface an agent is allowed to call, `frameworks/` and `harness/` for the runtime plumbing, `results/` for outcomes, `datasets/` for the published sets, `config/`, `scripts/`, `tests/` and `docs/` around them. The `traffic/` directory is the one name that does not announce its contents, and nothing in the visible page explains what it holds. A separate Chinese README sits beside the English one and is linked from it.

## Two more harnesses are vendored beside the main one

The benchmark runs inside the OpenClaw runtime, and version v2.0.3 records that the IronClaw and NanoClaw harness sources were vendored into the repository to support cross-harness evaluation. Both appear as top-level directories, `ironclaw/` and `nanoclaw/`, which means they are copies under version control rather than adapters resolved at install time. That choice has a purpose: a score that only holds on one harness measures that harness as much as the model. It also has a cost, since a vendored copy stops receiving upstream fixes unless someone merges them by hand, and the changelog shows the pattern of that work, with entries about syncing harness code updates and syncing benchmark bug fixes from the latest harness line. The `mock_tools/` directory fits the same approach, since a deterministic grade is easier to defend when the tools an agent calls are controlled rather than live.

## Deterministic grading, and the checkers that keep it that way

Two properties are claimed for the grading: determinism, and reliability across repeated trials. The changelog is where the work on them shows. Version v2.0.3 hardened multiple custom checkers with deterministic order-insensitive matchers, so two runs that produce the same content in a different order grade the same. The same entry records a resume model-alias preservation fix, which matters because a resumed run that silently changes the model under test is the easiest way to produce a number that cannot be reproduced. Earlier entries add scenario-grading fixes, trace-argument compatibility fixes for custom scoring, and `--exclude-scenario` filtering so a broken scenario can be taken out without editing the suite. The headline metric arrived in v1.0.9 as FinalScore, built on `pass^3`, `pass@3` and `average_score`, so a single number there already mixes repeated-trial consistency with breadth across trials and with an average, which is worth remembering before comparing it to a leaderboard that reports one pass rate.

## The open dataset is public, and the changelog says why that is a problem

The candidest paragraph in the whole page sits in the v1.1.4 entry. It adds two leaderboard models and then says that because the benchmark is fully open source it cannot fully avoid vendors optimising specifically for the public benchmark, and that a leaderboard based on a closed dataset will be released soon. That prediction is then kept: v2.0.0 launched the closed-dataset leaderboard with 33 model results, clickable model detail pages, visualisation charts and task browsing, and v2.0.4 published the open side, with the full 102 active scenarios on Hugging Face including prompts, fixtures and scenario configurations. So there are two datasets by design, and the paper title names the difference, trace-aware evaluation with runtime coverage and frozen workplace-style holdouts. The consequence for a reader is simple. An open score tells you what a model does on tasks anyone can read, and a closed score tells you what it does on tasks nobody has, and the two will not agree.

## Model access is arranged by the maintainer, not by the runner

The acknowledgements section is unusually concrete about how the leaderboard gets filled. It thanks friends from Kimi and Qwen for feedback on the leaderboard, and then thanks LongCat, Kimi, Ant Ling and MiMo for model access support, trial access or platform resources, stating that this reduced evaluation costs and made it possible to cover more frontier and preview models in a transparent live-runtime setting. Then the part that matters for reading a score: third-party API gateway providers whose served models are not run first-party, with Claude 4.7 Opus and GPT-5.5 given as examples, are invited to make contact, and the promise is that the benchmark can be run and reproducible results published when the evaluation setup is stable. Two consequences follow. A model on the board may have reached the harness through someone else's gateway rather than its own endpoint. And a result is published conditionally on stability, so the leaderboard is not a complete census of what exists, it is a list of what could be evaluated. The contact address is a single mailbox, xyh920691910@outlook.com, and it is also the address for submitting results.

## A second board keeps the disagreements visible

LLMLeadBoard sits in the same leaderboard family and does something different: it aggregates benchmark scores by model and inference mode from public model reports rather than from running anything. Its stated coverage is 63 models, 55 source reports and 444 benchmarks. The interface detail is the interesting part, because it handles disagreement instead of resolving it. Stacked cells appear when different reports give different scores for the same model and the same benchmark, so the reader sees that two sources disagree rather than being handed one of them. There are trend charts for benchmark progress inside a model family and iteration-speed charts for release cadence. Alongside the boards sit four blog posts, on the closed dataset release and its analysis, on safety under live agent work, on the author's own reflections during development, and on why the project was open sourced, all dated between April and May 2026. The paper is on arXiv and the dataset on Hugging Face, and the changelog runs from v1.0.7 to v2.0.4 while the repository itself publishes no GitHub releases at all.

## Conclusion

ClawProBench fits a team deciding between models on agent work rather than on chat quality, and its most useful feature is the refusal to collapse disagreement: the open dataset, the closed dataset and the report-sourced board are kept apart because each answers a different question. It does not fit anyone who wants a score to drop into a slide, since a scenario run needs the runtime, and the leaderboard entries depend on model access the maintainer arranges rather than on an API key you hold. Four things to check before quoting a number: which dataset it came from, since the open one is public and its own changelog says that invites targeted optimisation; which harness ran it, because the main one is OpenClaw and IronClaw and NanoClaw are vendored alongside for cross-harness comparison; whether the setup was called stable, which is the condition the page attaches to publishing results; and what FinalScore is made of, `pass^3`, `pass@3` and `average_score`. The paper is on arXiv as 2608.22510 and the dataset on Hugging Face. The repository publishes no GitHub releases despite a changelog running to v2.0.4, and the last commit is dated August 25, 2026.

## FAQ

### What does ClawProBench measure?

It evaluates LLM agents inside the OpenClaw runtime instead of in a static prompt setting, with deterministic grading, structured reports and benchmark-profile selection. The worktree inventory reports 102 active scenarios out of 164 catalog scenarios, 62 of them incubating.

### How do I list the ClawProBench scenarios?

With `python3 run.py inventory --json`, or `python3 run.py inventory --benchmark-status all --json` to include the incubating ones. The `core` profile is the default ranking path, with `intelligence`, `coverage`, `native` and `full` offering broader active coverage.

### Is the ClawProBench dataset public?

Both halves are. The open portion is on Hugging Face as xyh110sym/clawprobench, with v2.0.4 releasing prompts, fixtures and scenario configurations for all 102 active scenarios, and a separate closed dataset has its own leaderboard, which v2.0.0 launched with 33 model results.

### What is FinalScore in ClawProBench?

It is the metric introduced in v1.0.9, built on `pass^3`, `pass@3` and `average_score`. The harness claims two grading properties, deterministic grading and reliability across repeated trials, so one number there mixes consistency with breadth.

### How does ClawProBench get model access for its leaderboard?

From providers, not from the runner. The page thanks LongCat, Kimi, Ant Ling and MiMo for model access support, trial access or platform resources, and invites third-party API gateway providers whose served models it does not run first-party to get in touch, promising reproducible results when the evaluation setup is stable.

## Sources

- [Issues](https://github.com/suyoumo/ClawProBench/issues)
- [License: Apache-2.0](https://github.com/suyoumo/ClawProBench/blob/main/LICENSE)
- [Project website](https://suyoumo.github.io/bench/)
- [README](https://github.com/suyoumo/ClawProBench/blob/main/README.md)
- [suyoumo/ClawProBench on GitHub](https://github.com/suyoumo/ClawProBench)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/suyoumo-clawprobench
