# SkillsBench: a benchmark that only counts tasks the model fails

> SkillsBench measures whether a skill helps, and it goes about that by picking tasks deliberately designed so that strong models score under half. That is the inverse of the usual benchmark instinct, and it comes with an apparatus: a reference solution that must pass before any agent runs, a verifier that checks outputs, a sandbox per task, and one flag that turns the skill on.

**benchflow-ai/skillsbench** — SkillsBench evaluates how well skills work and how effective agents are at using them.

- Repository: https://github.com/benchflow-ai/skillsbench
- Website: https://www.skillsbench.ai
- Stars: 1,826 · Forks: 370
- Language: PDDL
- License: Apache-2.0
- Published: 2026-09-30 · Updated: 2026-09-30 · Language: en
- Canonical page: https://hysenlabs.com/projects/benchflow-ai-skillsbench

## The oracle has to pass before anything else

The quick start contains one line that is the whole methodology:

```bash
bench eval run --tasks-dir tasks/offer-letter-generator --agent oracle --sandbox modal
```

That runs the oracle, not a model. The comment above it says it plainly: the oracle must pass before an agent runs. So every task in this benchmark ships with a reference solution, and the first thing the harness does with a new task is check that the reference solution satisfies the task's own verifier.

That step is what makes the numbers mean something. A verifier is code, and code has two failure modes. Too strict and it rejects the correct answer, so every agent scores zero and the task measures the verifier. Too loose and it accepts anything, so every agent scores full marks and the task measures nothing. Neither failure is visible without a known-good solution to calibrate against, and the oracle is the calibration.

The task structure shows what is being checked:

```text
tasks/<task-id>/
  task.md
  environment/
    Dockerfile
    skills/
  oracle/
    solve.sh
  verifier/
    test.sh
    test_outputs.py
```

Four parts, each with a distinct job. The task statement is a document. The environment is a container that also holds the skills under test, which is worth noting: the skill is inside the thing being measured, not bolted on from outside, so it has access to whatever the task gives it and nothing more. The oracle is a shell script, which keeps the reference solution in a language the task itself can be written in rather than in the harness. And the verifier is a shell entry point plus a Python test, which tells you the grading is on files and outputs rather than on prose.

That last point is a real constraint on what can be measured. Anything whose success is a matter of taste cannot be expressed as a test script, so every task in this benchmark has been shaped to have a checkable result.

## Tasks are selected for being hard

The stated goals are where this benchmark takes its distinctive position, and one of them is a constraint rather than an aspiration.

The first goal is breadth and quality. The second is the interesting one: tasks are to be designed so that they require skill composition, meaning two or more skills rather than one, and so that state-of-the-art performance on them is under fifty percent. The third is a list of target models.

Almost every benchmark in this space selects tasks because they are representative. This one selects them because the current frontier does not already solve them, and it goes further by requiring composition rather than a single capability. The reasoning is coherent. If a model can already do the task unaided, then a skill cannot improve the number, so the task distinguishes nothing. If a task needs one skill, then you are measuring whether the agent found a file and read it. Composition is the case where the skill's value is structural: the agent has to know that two capabilities exist and order them correctly.

The fifty percent threshold is doing the work of a difficulty filter. Anything the target models pass comfortably is excluded from consideration, and anything they cannot approach at all is not informative either, so the interesting band is the awkward middle.

This is a bet about how long a benchmark stays useful, and it is worth naming the trade-off explicitly. A benchmark designed to be hard ages quickly, because the target models move and the tasks they cannot solve are precisely the ones they will learn to solve. The project is built for that, with two releases in the first three days and a weekly community call, but the consequence for a user is that a score from six months ago says very little about today unless you know which tasks and which models it used.

## The skill being measured has a pinned definition

A benchmark about skills needs a definition of what a skill is, and that definition is not written here. It comes from another repository and it is pinned.

The project manifest has a source entry for a package called the skills reference, taken from a git repository at a specific commit, with a subdirectory rather than a branch or a tag. So the format under evaluation is a fixed revision of an external specification, not whatever the ecosystem happens to look like this week.

That is the right call for a measurement tool, and it is easy to underrate. If the thing being measured is a moving target, a score is only comparable with itself, and comparing two skills becomes impossible because you cannot tell whether one of them used the format better or the format moved underneath it. Pinning the commit means a score from a given run names a specific definition of the unit.

It also means the maintenance question is delegated. The people who own the format are not the people who own the benchmark, and when the format changes this benchmark has to move deliberately rather than automatically. That is a trade you make once and then live with.

The repository has its own view of the skill space as well. There is a taxonomy, in both a human-readable form and a machine-readable one, and a registry file. Those are the two artefacts a benchmark of this kind needs and most do not have: a way to describe what categories of skill exist, and a machine-readable list of what is in scope. The registry in particular is what lets results be aggregated across runs and across submissions rather than living in one operator's notes.

## The experiment is one flag

The measurement, stripped to its essence, is that the same task is run twice with one variable changed.

```bash
bench eval run \
  --tasks-dir tasks-extra/mhc-layer-impl \
  --agent claude-agent-acp \
  --model <model> \
  --skill-mode with-skill \
  --skills-dir tasks-extra/mhc-layer-impl/environment/skills/ \
  --sandbox modal
```

The skill mode flag is the experiment. The skills directory points at the skills that ship inside that task, so the skill under test is the one the task defines rather than something the operator supplies. Everything else, the task, the verifier, the sandbox, the agent, the model, stays fixed.

That is what makes the result attributable. If you compare two agents' scores, the difference might be the model. If you compare two sandboxes, it might be the environment. Holding everything else at one value and changing the skill presence is the only way to answer the question the project says it is trying to answer, which is how effectively agents use skills rather than how well agents perform.

The featured task shows why composition is the target. It is a GPU task about implementing and comparing manifold-constrained hyper-connections in a training codebase, and it bundles three separate skills: one to launch the accelerator workflow, one carrying the algorithm, one for the training harness. An agent that does not know all three exist cannot get started, and an agent that knows all three has to order them correctly. That is the case a single-skill task cannot produce.

It is also the case where the cost stops being theoretical, because one of those skills starts a cloud GPU job. Benchmarking agent behaviour is not free once the tasks need hardware.

## Default tasks need nothing, extra tasks need everything

The task tree is split in two, and the split is about what you have to hand over before you can run something.

The default directory holds runnable tasks that need no external credentials. Those are what you get after cloning and syncing, and they are what the quick start uses. The extra directory holds tasks that depend on credentials or that are incompatible with the integration runner, and including them takes a deliberate flag that suppresses the default exclusions.

That flag is the design decision worth noticing. The default is a task set that runs anywhere with no setup, which means a new contributor gets a working evaluation in a few commands. The extra set is opt-in with an explicit acknowledgement, so nobody discovers at run time that their account is about to pay for accelerator time.

Sandboxing follows the same pattern. There is a default cloud provider and a local option for a Docker-based run instead, and the credentials for the cloud path are two environment variables. The provider is configurable because a benchmark that can only run on one vendor's infrastructure cannot be reproduced by someone who does not have an account there.

The repository also carries the reproducibility split, which is unusual and correct. The repository's own tooling is installed from the committed lockfile, with a flag that refuses to proceed if the lockfile does not match. But the benchmark framework itself is installed as a tool at its latest version, on the reasoning given in the README: day-to-day task authoring and evaluation should track the framework, while the repository's own scripts should not move underneath you.

Most projects pick one of those two positions and get it wrong half the time. This one states which applies to which.

## A web service, an agent protocol, and a queryable result store

The dependency list explains the architecture, and it explains it quickly.

```toml
"a2a-sdk[http-server]==0.3.20",
"benchflow[sandbox-daytona]>=0.6.3,<0.7",
```

One of those is pinned to an exact version with the HTTP server extra, and the other is bounded to a minor range. The exact one is the protocol SDK that agents use to talk to the harness, so the project has decided that surface should not move. The bounded one is the framework the harness is built on, which has decided the opposite.

That asymmetry is the interesting part, and it is a good proxy for how a project thinks about its own dependencies: pin what your results depend on, float what your tooling depends on.

The presence of an HTTP server extra, together with the web framework and ASGI server in the same list, tells you the benchmark is a service. Agents run inside sandboxes and reach back out to a running harness over a protocol, rather than the harness shelling into each sandbox. That is the right shape for a benchmark, because it means the harness is reachable from places the harness cannot reach.

The development dependencies include a columnar database. Results from a benchmark are not a folder of logs, they are something you want to query across runs and models, and putting a query engine in the development group rather than bolting one on later is the kind of decision that is cheap in month one and expensive in year two.

The lint configuration is worth one line too, because the ignore list is annotated with reasons, including that line-length is left to the formatter and that bare except clauses are considered common in scripts. A project that explains its lint exceptions has decided that the next contributor can evaluate them.

## Where this benchmark will mislead you

Three limits, all consequences of the design rather than flaws in it.

First, a score here is not a measure of agent quality. The tasks are selected because the target models do not solve them, so a high score means a skill closed a specific gap on a specific hard task. It does not mean the agent is generally capable, and a skill that improves an agent on tasks the agent already solves will register as nothing here. That is acceptable if you are choosing between skills, which is what the benchmark is for, and misleading if you are choosing between agents, which it is not for.

Second, difficulty is relative to the models listed, and those models move. A benchmark that refreshes its task set to stay hard is doing the right thing for its own validity and making your old results harder to interpret. The taxonomy and the registry files are the tools for checking whether the set is the same one your earlier run used.

Third, a verifier that is a shell script and a Python test can only check what can be asserted about files and outputs. That is a wide surface, and it excludes a category of task where success is a matter of judgement, which means the benchmark is quietly biased toward mechanical work. Nothing in the documentation claims otherwise, but the bias is worth naming because the skills you most want to justify are often the ones for tasks nobody can write a test for.

On maintenance: the licence is Apache 2.0, the last push was on 2026-07-23, and the releases are two in three days, a benchmark snapshot and then a packaging change. The project version in the manifest is still at 0.1.0, which does not match the release tags and is the kind of mismatch worth raising rather than working around.

## Conclusion

Adopt SkillsBench if you are deciding whether to ship a skill, because the with-skill and without-skill comparison on one task with one verifier is a shape of evidence that is hard to argue with. Do not read a score as a measure of agent quality: the tasks are selected for being beyond the targeted models, so a high score means the skill closed a specific gap rather than that the agent is capable. Check the pinned revision of the skills reference before comparing results across months, since the definition of a skill comes from another repository, and budget the sandbox cost before running the credentialed tasks.

## FAQ

### What does SkillsBench measure?

It measures how effectively agents use skills, which it defines as modular folders of instructions, scripts and resources for specialised workflows. Both skill effectiveness and agent behaviour are evaluated through gym-style benchmarking, running the same task with and without the skill available.

### What is an oracle in SkillsBench and why must it pass first?

Each task ships a reference solution under an oracle directory as a shell script. The harness runs it against the task's verifier before any agent runs, because a verifier that is too strict rejects correct answers and one that is too loose accepts anything. The oracle is what tells those two failure modes apart.

### How does SkillsBench pick its tasks?

The stated goal is tasks requiring composition of two or more skills where state-of-the-art model performance is below fifty percent. The reasoning is that a task the frontier already solves cannot show a skill's value, so the interesting band is tasks where the baseline is insufficient but not impossible.

### How do I run SkillsBench locally?

Clone the repository, install the benchmark framework as a tool, sync the repository tooling from its committed lockfile, then run a task. Modal is the default cloud sandbox and needs two token environment variables; pass a Docker sandbox option instead for a local-only run. Default tasks under `tasks/` require no external credentials.

### How are skills supplied to a SkillsBench task?

They live inside the task's own environment directory alongside its Dockerfile, and a run points the skills directory flag at that location. The skill under test is therefore the one the task defines rather than something the operator supplies, which is what makes two runs comparable.

### What is the structure of a SkillsBench task?

Four parts: a task statement document, an environment directory holding a Dockerfile and the skills, an oracle directory with a reference solution script, and a verifier directory with a test entry point plus a Python test that checks outputs.

## Sources

- [benchflow-ai/skillsbench on GitHub](https://github.com/benchflow-ai/skillsbench)
- [License: Apache-2.0](https://github.com/benchflow-ai/skillsbench/blob/main/LICENSE)
- [Project website](https://www.skillsbench.ai)
- [README](https://github.com/benchflow-ai/skillsbench/blob/main/README.md)
- [Releases](https://github.com/benchflow-ai/skillsbench/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/benchflow-ai-skillsbench
