# ClawBench: a browser-agent benchmark built around daily online tasks

> ClawBench packages a task suite, a Python evaluation harness and published traces for testing web agents on ordinary online work. It is a research instrument, not a CI check, and the README leaves several operational questions open.

**TIGER-AI-Lab/ClawBench** — Open-source benchmark for browser AI agents on daily tasks.

- Repository: https://github.com/TIGER-AI-Lab/ClawBench
- Website: https://claw-bench.com
- Stars: 903 · Forks: 62
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/tiger-ai-lab-clawbench

## The gap ClawBench fills: everyday browser tasks, not toy pages

Most web-agent evaluation happens on two extremes. Either the task is synthetic, a form on a page the maintainers wrote, or it is an enterprise workflow that requires credentials nobody outside the company has. ClawBench sits in between. Its stated subject is real-world online tasks, and the repository ships a test-cases directory plus a models directory with a schema, which tells you the two things the project treats as configuration: what to run and which model runs it.

The intended audience is narrow. If you are building or comparing browser agents, you need a task set that is fixed, a runner that produces comparable output, and traces you can inspect when a run fails. ClawBench provides all three: the package exposes a TUI and a batch runner, and the project publishes trace datasets on Hugging Face alongside the task dataset. If you are shipping a product and want a smoke test, this is the wrong shape of tool. The task definitions live in the repository, not in your application, and the scoring path runs through a model you configure.

The academic framing is explicit. The README carries an EMNLP 2026 Findings badge and links an arXiv paper, and the repository includes CITATION.cff. That is a signal about how results are meant to be consumed: as a paper artifact with a leaderboard, not as a vendor-neutral certification.

## How the harness is wired: bundled config, model schema, separate runners

The pyproject.toml is the clearest map of the architecture. The wheel build force-includes four things into the package under clawbench/_bundled: a .env file, models/models.example.yaml, models/model.schema.json, and the entire test-cases directory. So a pip install does not just give you code. It gives you the task set, an example model configuration and its JSON schema, and an environment file, all inside the installed package.

That design choice has a consequence worth naming. Task definitions and model configs are versioned with the library, not fetched at runtime. Upgrading the package can change the task set under you, which matters if you are comparing scores across weeks.

The entry points split responsibilities rather than hiding them behind one command. clawbench opens a TUI. clawbench-run executes a single run, clawbench-batch executes many. clawbench-rescore re-evaluates existing output, and clawbench-analyze aggregates. Two more, clawbench-reproduce and the harbor and edgebench adapters, exist to map ClawBench tasks onto other harnesses. There is also a separate prorl pair, clawbench-prorl-submit and clawbench-prorl-mock, which suggests a submission path with a mock gateway for local testing.

Dependencies are deliberately thin: fpdf2, huggingface_hub, pyyaml, questionary and rich. No browser automation library appears in the dependency list, which means the agent side is expected to bring its own stack and connect through the run interface rather than being driven by ClawBench itself.

## Installing clawbench-eval and running a first task

The package name on PyPI is clawbench-eval, not clawbench, and pyproject.toml sets requires-python to >=3.11. The README links the PyPI badge, so installation is a normal pip install rather than a source checkout.

```bash
pip install clawbench-eval
```

After that, the clawbench command should be on your PATH. It opens the interactive TUI, which is the entry point the README points at for getting started.

```bash
clawbench
```

Because the wheel force-includes models/models.example.yaml and models/model.schema.json, you will have an example model configuration to copy and edit. The pyproject.toml shows the schema file is shipped precisely so the config can be validated, so edit a copy rather than the bundled original.

Point your model configuration at a provider you actually have access to, then run a single task through the single-run entry point before attempting a batch. The batch entry point is clawbench-batch, and rescoring existing output without re-running the agent is clawbench-rescore. The README does not document what a successful first run prints, so treat the first invocation as a configuration check rather than an expected result.

## Where ClawBench gets in your way

The bundled .env is the first obstacle. Shipping an environment file inside a wheel is unusual, and it means credentials and defaults are part of the installed artifact rather than something you supply from outside. Anyone running this in a shared environment should look at what that file contains before the package reaches a machine with real keys on it.

The second limitation is evaluation cost. Scoring a browser agent means driving a browser and making model calls, and ClawBench does not abstract that away. The dependency list has no browser library, so whatever automation you use is your own integration. A batch run over the full test set is therefore a real spend in both time and API budget, and the repository does not document a cheap subset mode.

The third is reproducibility across versions. Releases are frequent: v0.9.1 on 2026-08-08, v0.9.2 on 2026-08-19, v0.10.0 on 2026-08-30. With tasks bundled into the wheel, a pinned version is the only way to guarantee that two runs scored the same task set. The README does not document rollback or a compatibility policy between minor versions.

Finally, this is not a deployment gate. It measures whether an agent completes everyday online tasks under benchmark conditions. It says nothing about latency, cost per task, or behaviour under adversarial input, and the README makes no claims about any of those.

## ClawBench against WebArena-style self-hosted environments

The obvious comparison is with benchmarks that host their own websites. WebArena and similar projects stand up containerised replicas of real sites, so the environment is fully controlled and reproducible offline, at the cost of building and maintaining that infrastructure and accepting that the sites are approximations.

ClawBench takes the other route. Its subject is real-world online tasks, and it publishes traces on Hugging Face rather than shipping a simulated web. The trade is inverted: you get tasks that reflect actual sites and actual workflows, and in exchange you lose the guarantee that the page you hit today is the page the reference run hit. That is why the trace datasets matter so much here. They are the evidence that a run happened the way it is reported, and they are the only way to check a failure after the fact when the live page has moved on.

There is a second, quieter difference. WebArena-style setups ship the browser and the sites together, so the harness controls the whole stack. ClawBench ships the tasks and the scoring paths and leaves the browser to you. That makes it easier to plug in whatever agent you already have, and harder to guarantee that two people running the same task used the same tooling.

## Maintenance, licensing and what an upgrade actually costs

The repository is not archived, and the last push was on 2026-09-10. With three releases in the six weeks before that, the project is moving. That cuts both ways for adopters: fixes arrive quickly, and so do changes to bundled tasks.

The licence is Apache-2.0, and the repository carries a separate NOTICE file, which the wheel force-includes. Apache-2.0 permits commercial use and modification with attribution and notice retention. The NOTICE file exists because some bundled material may carry its own terms, and the test-cases directory is bundled wholesale into the wheel. If you redistribute clawbench-eval or a derivative, read NOTICE and the individual test case files rather than assuming the top-level licence covers everything. This is a description of the repository layout, not legal advice.

Upgrade cost is dominated by the packaging decision. Because models/models.example.yaml, models/model.schema.json and test-cases are all inside the wheel, a version bump can change your configuration surface and your task set at the same time. Pin the version in your environment, keep your edited model config outside the package, and re-run clawbench-rescore against stored output rather than re-running the agent when you only want to check a scoring change.

## Conclusion

Adopt ClawBench if you are comparing web agents or reproducing published numbers, and you accept that scoring depends on model credentials you supply yourself. Do not adopt it as a pass/fail gate in a deployment pipeline: it is a research benchmark with a fast release cadence and no documented rollback path. Before committing, check the models/models.example.yaml schema against your own provider, confirm which task subset you intend to run, and read the licence and NOTICE files for the bundled test cases and assets.

## FAQ

### What is ClawBench and what does it measure?

ClawBench is an open-source benchmark for browser AI agents on daily online tasks, distributed as the clawbench-eval Python package. It ships a task set, a model configuration schema and runners, and the project publishes task and trace datasets on Hugging Face alongside a leaderboard.

### How do I install ClawBench?

Install the clawbench-eval package with pip, which requires Python 3.11 or newer, then run the clawbench command to open the interactive TUI. The wheel bundles an example model configuration and the test cases inside the package.

### Where is the ClawBench dataset?

The README links Hugging Face datasets for the tasks and for the V2 traces, and a separate V1 traces dataset. The wheel also bundles the test-cases directory, so the task definitions ship with the installed package.

### Does ClawBench need model API credentials?

Yes. The package ships a models directory with an example configuration and a JSON schema, and the wheel force-includes a .env file, so you supply and validate your own model configuration before running tasks.

## Sources

- [License: Apache-2.0](https://github.com/TIGER-AI-Lab/ClawBench/blob/main/LICENSE)
- [Project website](https://claw-bench.com)
- [README](https://github.com/TIGER-AI-Lab/ClawBench/blob/main/README.md)
- [Releases](https://github.com/TIGER-AI-Lab/ClawBench/releases)
- [TIGER-AI-Lab/ClawBench on GitHub](https://github.com/TIGER-AI-Lab/ClawBench)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/tiger-ai-lab-clawbench
