WildClawBench: an in-the-wild agent benchmark that runs inside a live OpenClaw instance
An in-the-wild benchmark for AI agents in the production harness.
At a glance
- What is it?
- WildClawBench puts 60 hand-written tasks in front of four agent harnesses inside a real OpenClaw environment and grades them in isolated Docker containers. It is for people who need to tell model capability apart from harness scaffolding, and the README is candid that most models still score far below the top result.
- Who is it for?
- Adopt WildClawBench if you are comparing agent harnesses or frontier models on long-horizon, tool-heavy work and you can afford Docker plus an OpenRouter key. Do not adopt it as a quick smoke test for a small model, and do not treat the leaderboard as a substitute for running the suite yourself.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 12 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap WildClawBench was built to fill
Most agent evaluations isolate a single skill. Call this function, parse that JSON, follow one instruction. WildClawBench takes the opposite position: the README calls it "hard, practical, end-to-end evaluation for AI agents", and the task list backs that up. Clipping goal highlights from a football match, negotiating meeting times across multiple email rounds, hunting contradictions in search results, writing inference scripts for an undocumented codebase, catching privacy leaks before they happen.
The intended audience is narrow and specific. If you are choosing between agent harnesses, or between models that will run inside one, a benchmark that only measures the model in a bare loop tells you little. WildClawBench runs the same 60 tasks through OpenClaw, Claude Code, Codex CLI and Hermes Agent under one grading scheme. That design choice is the whole point: it separates what the model can do from what the surrounding scaffolding lets it do.
The README is also unusually direct about headroom. In the technical report snapshot the strongest frontier model reached 62.2% overall, and audited OpenClaw runs have since raised the best score to 67.2%. Most models land well below that. A benchmark where the ceiling is in the sixties is doing its job.
How a task actually executes: containers, injection, and grading
Each task runs in its own Docker container. The README states that the same image, the same data and the same grading code are used every time, and that ground truth and grading scripts are injected only after the agent finishes. They are never visible during execution. That ordering is the anti-leakage mechanism, and it is the detail worth checking before you trust a score.
The environment is a live OpenClaw instance with real tools attached: browser, bash, file system, email, calendar. Tasks are not stubs against a mock API. The README describes agents chaining 10 to 60 or more tool calls, adapting when services fail, and deciding what to do rather than only how. Wall-clock horizons of 10 to 20 minutes appear in the long-horizon category.
The task taxonomy in the README has five axes: agency, multimodal, long-horizon, coding and safety. Multimodal work includes tracking events across a 45-minute match video and classifying 12 clothing photos into 4 styled outfits with generated full-body images. Safety work buries harmful instructions deep inside normal-looking documents and scatters API keys across a large git history. Those are grading problems as much as agent problems, which is why the judge model matters.
Installing the evaluation pipeline and running a first task
The repository is Python and the evaluation code lives in this repo; the task data ships separately on Hugging Face. Start by cloning and installing the pinned requirements, which cover document rendering, HTTP, video download, ModelScope and dotenv handling.
pip install -r requirements.txtNext, copy the environment template. The file .env.example is the authoritative list of keys, and every value below is copied from it. Note that DOCKER_IMAGE is a tag you must have locally, and DEFAULT_MODEL is a placeholder you have to replace.
cp .env.example .envDOCKER_IMAGE=wildclawbench-ubuntu:v1.3
GATEWAY_PORT=18789
TMP_WORKSPACE=/tmp_workspace
TASKS_SUBDIR=tasks
OUTPUT_SUBDIR=output
DEFAULT_MODEL=openrouter/xxx
DEFAULT_PARALLEL=1
JUDGE_MODEL=openai/gpt-5.4
OPENROUTER_API_KEY=
BRAVE_API_KEY=The template also reserves keys for a custom endpoint (MY_PROXY_API_KEY) and for skills that need their own credentials (GEMINI_API_KEY, FIRECRAWL_API_KEY, EXA_API_KEY). Those are commented out, which tells you they are optional until a task requires them. The proxy variables HTTP_PROXY_INNER, HTTPS_PROXY_INNER and NO_PROXY_INNER apply inside the container, not on your host.
If you would rather not touch this pipeline at all, the README points to a second path: the WildClawBench-Harbor dataset repackages all 60 tasks in Harbor format so that a single harbor run executes them with no benchmark-specific setup. That is the lower-friction entry point, and it is the one to try first if your interest is the tasks rather than the harness comparison.
Where WildClawBench will disappoint you
The heaviest constraint is Docker. Every task runs in its own container against a specific image tag, and the README does not document a fallback for environments where nested containers or a Docker daemon are unavailable. If you cannot run Docker, this benchmark is not for you, and the Harbor route does not remove that dependency.
Grading depends on a judge model, configured through JUDGE_MODEL and OPENROUTER_API_KEY. That means your scores are partly a function of a model you chose and paid for, and the README does not describe a judge-validation or agreement procedure. Treat cross-run comparisons cautiously unless the judge configuration is held fixed.
Cost and time are real. Tasks span 10 to 20 minutes of wall-clock execution with 10 to 60+ tool calls, and DEFAULT_PARALLEL defaults to 1. A full 60-task sweep across four harnesses is not an afternoon's work on a laptop, and the README gives no cost estimate.
Finally, the safety category is the one most likely to produce results you cannot act on. Detecting a credential leak buried in git history or refusing an instruction hidden in a document is a genuine capability, but a single benchmark score does not tell you how the agent behaves on your own documents. Use it as a signal, not a certification.
How it differs from Harbor-format and single-harness benchmarks
The closest thing to an alternative inside this project's own ecosystem is the Harbor format itself. WildClawBench-Harbor repackages the same 60 tasks so any Harbor-supported agent can run them with one command. The difference in approach is packaging versus pipeline: Harbor gives you portability across agents at the cost of the four-harness comparison the main repository is built around. If your question is "how does my agent do on these tasks", Harbor is enough. If your question is "how much of the score comes from the harness", you need the repository's own pipeline, because that is where the same suite is run under OpenClaw, Claude Code, Codex CLI and Hermes Agent with one grading scheme.
The other distinction is the environment. Benchmarks that mock tool APIs are cheaper to run and easier to reproduce, and they answer a narrower question. WildClawBench trades that convenience for a live OpenClaw instance with browser, bash, file system, email and calendar attached, plus failure modes that come from services actually breaking. The README's own framing, "real environment, not mocks", is the trade stated plainly: more realism, more moving parts, more ways for a run to fail for reasons unrelated to the model.
Licence, maintenance and the cost of staying current
The repository is MIT licensed, which is permissive for the evaluation code. The benchmark data is a separate matter: task data, Docker images and trajectories live on Hugging Face as three datasets, and the licence terms of those datasets are not stated in the README. Check them separately before redistributing task content, and note that the README points to a technical report PDF and an arXiv entry for the methodology. None of this is legal advice; read the actual licence files.
On maintenance, the last push to the repository was on 2026-08-17, which is recent. The README's news section shows an active release cadence through 2026: four harnesses added in May, frontier-model leaderboard expansions in July, and the Harbor and Trajectories datasets in August. Two external releases, Meta's Muse Glimmer and ByteDance Seed's Seed2.1, report WildClawBench scores, which is a form of adoption the README can point to.
Upgrade cost is the part to plan for. The Docker image is versioned (wildclawbench-ubuntu:v1.3), the harness roster has already grown once, and the leaderboard is described as continuously updated as new models are evaluated. Pinning the image tag and recording the judge model alongside any score you publish is the minimum discipline for making your numbers comparable six months from now.
Editorial conclusion
Adopt WildClawBench if you are comparing agent harnesses or frontier models on long-horizon, tool-heavy work and you can afford Docker plus an OpenRouter key. Do not adopt it as a quick smoke test for a small model, and do not treat the leaderboard as a substitute for running the suite yourself. Before trusting a number, verify the DOCKER_IMAGE tag in .env.example matches the image you built, confirm JUDGE_MODEL is set to a model you are willing to pay for, and check that the task directory you point TASKS_SUBDIR at is the one you intend to grade.
Frequently asked questions
What is WildClawBench?
It is an agent benchmark with 60 original tasks that run end-to-end inside a live OpenClaw environment, with browser, bash, file system, email and calendar tools attached. The same task suite is executed under four harnesses: OpenClaw, Claude Code, Codex CLI and Hermes Agent.
How do I install WildClawBench and run it?
Clone the repository, run pip install -r requirements.txt, then copy .env.example to .env and fill in the values. The template expects a local Docker image tag such as wildclawbench-ubuntu:v1.3, a model for DEFAULT_MODEL, and OPENROUTER_API_KEY plus JUDGE_MODEL for grading.
Can I run WildClawBench without this repository's pipeline?
Yes. The README says WildClawBench-Harbor repackages all 60 tasks in Harbor format so any Harbor-supported agent can run them with a single harbor run and no benchmark-specific setup.
Does WildClawBench require Docker?
Yes. The README states that each task runs in its own Docker container with the same image, data and grading code, and .env.example sets DOCKER_IMAGE to a specific image tag. The README does not document a non-Docker path.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/internlm-wildclawbench)