Model or dataset
pinchbench/skill avatar
pinchbench/skill

PinchBench: the task count is machine-written, the judge runs inside the harness, and the dates have gone quiet

PinchBench is a benchmarking system for evaluating LLM models as OpenClaw coding agents. Made with 🦀 by the humans at https://kilo.ai

1,356 stars159 forksPythonMIT

At a glance

What is it?
A benchmark that runs LLM models as the brain of an OpenClaw agent across 53 real-world tasks. The interesting mechanics are the grading path, the flag matrix for suite and thinking level, and the fact that the official leaderboard is driven from a separate repository.
Who is it for?
PinchBench suits someone who already runs OpenClaw and wants to know how a model behaves as the agent's brain rather than as a chatbot, and who is willing to read the flag matrix before trusting a score. Two caveats carry the weight.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 93 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The task count is injected, not typed

The number 53 does not sit in the prose as a fact anyone maintains by hand. It is wrapped between a pair of HTML comment markers, a text variant reading PinchBench includes 53 tasks across real-world categories, and a badge variant above it. That means the count is emitted by tooling into the README, and what you read is a snapshot of whatever the tooling last produced. If the tasks directory changes without the injection being re-run, the headline number drifts. The eight categories underneath are ordinary markdown and list the shape of the suite: Productivity, Research, Writing, Coding, Analysis, Email, Memory and Skills, with the last covering ClawHub and skill discovery inside the OpenClaw ecosystem.

The judge runs inside the harness by default

This is the methodological detail that decides whether a score means anything. With no --judge flag, the LLM judge runs as an OpenClaw agent session, so grading happens inside the same framework the candidate model is being tested in. Set --judge and it calls the model API directly instead, which the documentation describes as bypassing OpenClaw personality injection. Each task is graded automatically, by an LLM judge, or both, so a run can mix deterministic checks with model judgement. Four direct routes are shown, OpenRouter, the Kilo Gateway, Anthropic and OpenAI, plus a headless Claude CLI variant whose command line is visible only as far as --model open. Which model is the default judge is not stated anywhere, and no --judge value means it is whatever the session falls back to.

Eleven flags, a suite selector that takes task IDs

The command reference is where the real surface is. --suite takes all, automated-only, or a comma-separated list of task IDs, so a run can be narrowed to task_calendar,task_stock rather than filtered by category name. --runs N repeats each task for averaging, --timeout-multiplier N scales timeouts for slower models, and --thinking LEVEL takes off, minimal, low, medium, high, xhigh or adaptive. Output lands in results/ by default and can be redirected with --output-dir. Submission is a separate concern from running: --register requests a one-time API token, --no-upload keeps results local, --upload FILE pushes a previous results JSON, and --official-key or PINCHBENCH_OFFICIAL_KEY is what marks a run as official on the board. One script covers both halves.

The official leaderboard is driven from a different repository

The opening note is unambiguous about what this repository is not. It contains the benchmark skill and the tasks, and it is not the source of official leaderboard results. To add a model to the official results you modify a file called default-models.yml, and the link points at github.com/pinchbench/scripts rather than at this repository. So a project named PinchBench, published as pinchbench-skill, drives its official board from a sibling repository that is not in the tree here. The practical consequence for anyone reading scores: the model roster on the board and the code being cloned are maintained separately, and nothing in the run output verifies that a given model was on the official list at the time it was measured.

A running OpenClaw instance is a hard prerequisite

The runner is one shell script. To clone the skill and run the suite:

bash
# Clone the skill
git clone https://github.com/pinchbench/skill.git
cd skill

# Run benchmarks with your model of choice
./scripts/run.sh --model openrouter/anthropic/claude-sonnet-4

# Or run specific tasks
./scripts/run.sh --model openrouter/openai/gpt-4o --suite task_calendar,task_stock

Requirements are listed as Python 3.10+, the uv package manager, and a running OpenClaw instance. That third item is the one that shapes how you can use this. The benchmark measures a model as the brain of an OpenClaw agent, so there has to be an agent runtime already present and reachable before any of the 53 tasks can run; this is not a standalone harness you point at an API key. A live demo at pinchbench.com exists separately from the local runner. Model identifiers must carry their provider prefix, with examples such as openrouter/ and anthropic/, and OpenRouter is named as the default provider used for routing. Getting the prefix wrong is the most likely reason a first run fails to route at all.

Dependencies are three runtime packages and an SSH library

The manifest is short enough to read in one go. Runtime dependencies are pyyaml, fabric and paramiko, with pyyaml at 6.0.1 or newer and the other two pinned to a minor floor. Fabric and paramiko together are the surprise: paramiko is an SSH client for Python, and fabric builds on it, so the benchmark's install footprint includes SSH capability that nothing in the task list obviously needs. Dev extras add pytest, pytest-cov, black and ruff. Both black and ruff are configured to line-length 100 targeting py310, so the two formatters and linters agree on width. Versioning is dynamic through setuptools-scm with no hardcoded number, and a console script named benchmark points at benchmark:main.

Crab, SKILL.md, skills-lock.json and a benchmark Dockerfile

The tree mixes an agent skill with a Python project and does not pretend otherwise. SKILL.md sits at the root and is the file the packaging metadata names as the readme, so the distribution's long description is the skill definition rather than README.md, which is why the two documents can drift apart in tone. skills-lock.json suggests the skill is installed through a lockfile of some kind, and a .agents directory sits alongside .github and a pre-commit config. There is also crab.txt, a single named file with no counterpart in the README, and a Dockerfile.benchmark distinct from the skill itself, implying the benchmark runs in a container rather than directly against the host's OpenClaw install.

Dated tags, then three months of quiet

Maintenance state is worth stating plainly. The last push to the repository is dated 2026-07-02, and the newest tag is v2.0.0 from 2026-05-06, preceded on the same afternoon by v2.0.0-rc9 and v2.0.0-rc8 on 2026-04-27. That is two release candidates and a final inside about nine days, then a two-month gap to the last push, and three months of nothing since. A BENCHMARK_VERSION file at the root suggests the benchmark carries its own version marker separate from the package version. Nothing here is marked archived. If you are weighing whether the task suite still tracks current models, the gap between the last push and today is the fact to weigh, not the tag list.

Editorial conclusion

PinchBench suits someone who already runs OpenClaw and wants to know how a model behaves as the agent's brain rather than as a chatbot, and who is willing to read the flag matrix before trusting a score. Two caveats carry the weight. The default grading path puts the judge inside the same OpenClaw session as the candidate, so pass --judge to get a direct API call and remove the harness from the scoring side. And the official leaderboard is not produced by this repository; its model list lives in a separate scripts repository, so a green local run is a local run until you register and submit.

Frequently asked questions

What does PinchBench actually measure?

It measures how well LLM models perform as the brain of an OpenClaw agent, using 53 real-world tasks across eight categories rather than synthetic tests: tool usage, multi-step reasoning, ambiguous instructions, and whether the practical outcome happened.

How does the PinchBench LLM judge work by default?

With no --judge flag the judge runs as an OpenClaw agent session, so grading happens inside the same framework the candidate is tested in. Supplying --judge calls the model API directly instead, bypassing OpenClaw personality injection.

What do I need before I can run PinchBench tasks?

Python 3.10 or newer, the uv package manager, and a running OpenClaw instance. Model IDs must carry their provider prefix, such as openrouter/ or anthropic/, with OpenRouter as the default routing provider.

How do I get a PinchBench run onto the public leaderboard?

Run ./scripts/run.sh --register once to request an API token, then run the benchmark so results auto-upload with it. Use --no-upload to keep results local, and an official key, by flag or by PINCHBENCH_OFFICIAL_KEY, to mark a run as official.

Does the PinchBench repository produce the official leaderboard?

No. It holds the benchmark skill and the tasks, and the note at the top says it is not the source of official leaderboard results. Adding a model to official results means editing default-models.yml in the separate pinchbench/scripts repository.

Official sources

  1. License: MIT
  2. pinchbench/skill on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/pinchbench-skill.svg)](https://hysenlabs.com/projects/pinchbench-skill)