Model or dataset
claw-eval/claw-eval avatar
claw-eval/claw-eval

claw-eval: a benchmark whose grader is a different model

Claw-Eval is an evaluation harness for evaluating LLM as agents. All tasks verified by humans.

778 stars76 forksPythonMIT

At a glance

What is it?
claw-eval/claw-eval is an evaluation harness of 300 human-verified tasks with 2,159 rubrics across three splits, where the primary metric requires a model to pass three independent trials, the judge for the conversational split plays the simulated user as well as grading it, reproducibility is stated as still under audit, and the live web access the runs depend on is supplied by a paid sponsor.
Who is it for?
claw-eval is the most interesting of the agent benchmarks because it is specific about its own weaknesses. It requires three clean passes rather than a best-of-three, it names which model grades, it admits the codebase is still being audited so anyone can reproduce the leaderboard, and it adapts its tasks from six named public benchmarks rather than claiming a closed world.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Pass^3 is a conjunction, not a best of three

The primary metric was updated in March 2026 and the change is the point. To eliminate what the documentation calls lucky runs, a model has to pass a task across three independent trials before it earns a success credit, and the strict criterion says a task is marked as passed only if the success criteria are met in all three runs. So a model that solves a task twice out of three scores zero for it, which is a much higher bar than averaging or taking the best. The same logic is why robustness is one of the three grading dimensions rather than an afterthought: consistency across trials is the thing being measured. The dimension list is completion, whether the agent finished; safety, whether it avoided harmful or unauthorised actions; and robustness, whether it passes consistently. All three are assessed through what the documentation calls full-trajectory auditing rather than by looking only at the final answer.

API instability is handled by re-triggering the run by hand

Two lines in the methodology section are worth reading twice. The first is about reproducibility, and it is stated in the present tense as unfinished: the project is committed to end-to-end reproducibility, and the codebase is currently being audited to ensure all benchmark results on the leaderboard can be verified by the community. The second is about what happens when a run breaks for reasons that have nothing to do with the model. If execution errors are caused by network or API fluctuations, the evaluation is manually re-triggered so that exactly three trajectories are successfully generated. That is a defensible choice, since discarding a task because a provider had a bad minute would bias the leaderboard, but manual re-triggering is a human in the loop and it is not the same thing as an automated retry policy with a documented rule. Both of these are the honest place to look when deciding how much weight a leaderboard position carries.

In one split the grader also plays the user

A note beside the quick start names the grading models outright. Gemini 3 Flash is used for the general and multimodal tasks, while Claude Opus 4.6 is used for both the grader and the user agent in the multi-turn tasks. The first arrangement is the normal and defensible one, since a judge model that is not the model under test at least removes self-preference as a confound. The second is not the same thing, because in the conversational split the grading model is also the simulated user asking the questions. A model that sustains a conversation well with itself is being judged by itself in the role of the other party, and the 38 tasks in that split are the ones about clarification and advice, where the quality of the exchange is exactly what is under assessment. It is disclosed plainly, which is more than most harnesses do, and it is the first thing to check before comparing a multi-turn score with a general one.

requirements.txt says it matches two extras and does not

The dependency file opens with a comment claiming an equivalence: it says it is equivalent to installing the package in editable mode with the mock and sandbox extras. Comparing the two lists, the equivalence does not hold. The mock extra declares five packages: a web framework, an ASGI server, a PDF library, an HTML extraction library, and an HTTP client. `requirements.txt` lists the first three and then jumps to the Docker package for sandbox support. The HTML extraction and HTTP client packages are missing, which means anyone following the comment rather than the extras installs less than they were told they were installing, and the symptom would appear at the point where a task needs a page fetched or an article parsed. It is a small documentation defect in a file whose whole purpose is to be the reference.

The run command points at a directory the repository does not have

The batch invocation names a model configuration by path:

bash
claw-eval batch --config model_configs/claude_opus_46.yaml --sandbox --trials 3 --parallel 16

There is no `model_configs` directory in the repository. What is at the root is three loose YAML files, one for general tasks, one for multimodal, and one for the user agent, which the same line tells you to use for different task sets. So the documented example config path and the actual layout disagree, and the only way to run the harness is to construct the directory or edit the command. The same kind of mismatch appears in the dataset schema, where the fixture field is documented as pointing into a fixtures archive under a data directory that is not in the tree, with a separate note explaining that the complete fixtures including all video are on Hugging Face because of file size limits. The tasks themselves do live in the repository, under a tasks directory.

The live web the benchmark needs comes from a paid sponsor

Two keys are exported before a run, and both are paid. One is an OpenRouter key for model access. The other is a search key, and the comment beside it says to add it for tasks that need real web search, with a convenience link to a vendor. There is a whole section given over to that same vendor, describing a web scraper API, a web unblocker, and residential proxies provided for live-web agent tasks, framed as a way to keep web-dependent evaluations reproducible. A benchmark whose web-dependent tasks depend on a third party's proxy pool has a reproducibility story that runs through someone else's infrastructure, and the link in the section carries a referral parameter. That is not a criticism of the arrangement, which is how open benchmarks usually pay for themselves, but it belongs in your reading of any score that came from a task involving a live page.

The manifest says 1.0.0 and the newest documented version is 1.1.0

The package metadata declares version `1.0.0` while the most recent update entry in the documentation is v1.1.0, described as the release with 300 human-verified tasks across nine categories. There are no GitHub releases, so the version history exists only as three bullet points in the README: v1.1.0 for the current task set, v1.0.0 for building on reproducible real-world complexity, and v0.0.0 for going from chatbot to the real world, dated March 2026. A reader who checks the manifest and a reader who reads the changelog will therefore disagree about which version they are looking at, and there is no tag to arbitrate. The same pattern appears in the file layout, where three task configuration files sit loose at the repository root next to two cleanup and summary scripts, a sandbox requirements file, and a committed screenshot.

Joining the leaderboard is an email conversation

To run the harness and submit results you contact one of three addresses. There is no submission form, no automated upload path, and no documented review process in the documentation, which means the leaderboard is assembled by hand from results people send in. That has an upside, since a human reads every entry, and a downside, since the comparison depends on teams running the same protocol on the same day with the same provider. The acknowledgements are unusually concrete about where the tasks came from: the test cases are built on community work and adapted from tasks contributed by OpenClaw, PinchBench, OfficeQA, OneMillion-Bench, Finance Agent, and Terminal-Bench 2.0. Citing your sources is a good sign for a benchmark. Eight core contributors are listed, most from one university with two from another, alongside three advisors.

Editorial conclusion

claw-eval is the most interesting of the agent benchmarks because it is specific about its own weaknesses. It requires three clean passes rather than a best-of-three, it names which model grades, it admits the codebase is still being audited so anyone can reproduce the leaderboard, and it adapts its tasks from six named public benchmarks rather than claiming a closed world. Two things to weigh before you quote a number from it. The judge is not the same model as the one under test, and in the conversational split the judge also plays the simulated user, which means a model good at sustained conversation with itself scores well by construction. And the runs need a paid search key and a paid scraping proxy, so a result depends on a third party's infrastructure as well as on the model. Read the leaderboard's stated methodology before you treat a score as a measurement.

Frequently asked questions

How do you evaluate LLM agents?

With verified tasks, rubrics, and a pass criterion that survives unlucky runs. Claw-Eval uses 300 human-verified tasks with 2,159 rubrics across nine categories, grades on completion, safety, and robustness through full-trajectory auditing, and counts a task as passed only when the model meets the criteria in all three trials.

What is Claw-Eval?

An evaluation harness for evaluating language models as agents. It ships 300 human-verified tasks with rubrics across three splits, grades trajectories on completion, safety, and robustness, and publishes a leaderboard on its project site with a paper on arXiv.

How many tasks does Claw-Eval contain and how are they split?

300 tasks in total, each with human-verified rubrics, 2,159 rubrics altogether. The general split holds 161 tasks, multimodal holds 101, and multi_turn holds 38, the last being conversational tasks with simulated user personas for clarification and advice.

Which models does Claw-Eval use to grade agents?

Gemini 3 Flash grades the general and multimodal tasks. Claude Opus 4.6 grades the multi-turn tasks, where it also acts as the simulated user agent, which is worth knowing before comparing scores across splits.

Official sources

  1. claw-eval/claw-eval on GitHub
  2. Issues
  3. License: MIT
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/claw-eval-claw-eval.svg)](https://hysenlabs.com/projects/claw-eval-claw-eval)