# skill-up: a CLI that turns Agent Skills evaluation into a repeatable loop

> Alibaba's skill-up evaluates Agent Skills with declarative YAML cases across multiple agent engines, then hands the failures to skill-upper so an agent can repair the evals and rerun them. Here is how the two halves fit together, and where the tooling stops.

**alibaba/skill-up** — An evaluation and evolution tool for Agent Skills.

- Repository: https://github.com/alibaba/skill-up
- Website: https://alibaba.github.io/skill-up/
- Stars: 1,118 · Forks: 94
- Language: Go
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/alibaba-skill-up

## The gap skill-up fills between writing a SKILL.md and knowing it works

An Agent Skill is mostly prose. A SKILL.md describes when the skill should trigger, what steps it performs and what output it produces, and the agent decides at runtime how much of that to follow. That makes review awkward: two reviewers can read the same file and disagree about whether it works, and a change that fixes one behaviour can quietly break another.

skill-up targets that gap. The README describes it as an evaluation and evolution tool, and the evaluation half is the part that produces evidence: declarative YAML cases run across multiple Agent Engines, judged by rules, scripts or an agent, with results written to structured reports. The project positions itself against the workflow in the official Agent Skills evaluation guide, which it summarises as writing realistic cases, running with and without the skill, grading outputs, aggregating and iterating. skill-up's claim is that it replaces ad hoc run folders with a declarative format and automates workspace setup, skill installation, engine invocation, judging and report generation.

The audience is narrow and worth stating plainly. This is for people who maintain a skill and want a regression suite around it, and for teams that want that suite to run in CI. It is not for someone who has not written a skill yet.

## How eval.yaml, cases and engines fit together

The configuration is split in two. An eval.yaml defines the evaluation environment, the engine, the model and the case set; the cases themselves live in cases/*.yaml. Judging is pluggable, with rule_based, script and agent_judge listed as the supported strategies, and the report layer emits grading.json and benchmark.json in an Anthropic-compatible shape, plus benchmark.md, result.json, JUnit XML and HTML.

Engines are the execution half. Qoder CLI, Claude Code and Codex are built in, and the README also lists qwen_code among the supported engines. A user-defined agent goes through engine.custom, which the README describes as using a local transport and points at docs/design/custom-engine.md for the detail. That is the extension point to read first if your agent is not one of the built-ins, because everything else in the pipeline assumes an engine can be invoked and its output captured.

The evolution half is separate software. skill-upper is an Agent Skill shipped in the repository under skills/skill-upper. It reads failures from a report, decides whether the skill or the eval case is wrong, repairs or expands the suite, and reruns skill-up. The loop is conversational rather than a daemon: you keep talking to your agent and it keeps editing files and invoking the CLI.

## Installing skill-upper and running a first evaluation

The README recommends starting with skill-upper rather than the CLI. It installs through the skills package manager, and the command differs per agent. For Codex, a global install:

```bash
npx skills add https://github.com/alibaba/skill-up/tree/main/skills/skill-upper -g -a codex -y
```

For Claude Code, the same command with a different agent flag:

```bash
npx skills add https://github.com/alibaba/skill-up/tree/main/skills/skill-upper -g -a claude-code -y
```

According to the README you normally do not need to install skill-up first, because skill-upper checks for the CLI when it runs and walks the agent through installation if it is missing. If you would rather install the CLI yourself, the README gives a shell installer:

```bash
curl -fsSL https://raw.githubusercontent.com/alibaba/skill-up/main/install.sh | bash
```

With skill-upper installed, open a project containing your skill's SKILL.md in a compatible agent and ask it to evaluate the skill, naming the behaviours that matter and asking for realistic cases, validation and a run. The README's example prompt asks the agent to read SKILL.md, identify its most important behaviours, create cases with appropriate judges, validate the configuration, run skill-up and summarise the highest-impact failures. The layout it produces puts eval.yaml and cases/ under an evals directory next to SKILL.md, with run output in a separate workspace directory, for example my-skill-workspace/iteration-1/result.json. To iterate, you continue the same conversation and ask the agent to decide, per failure, whether the skill or the eval is wrong, then fix, add regression cases and rerun.

## The evolution loop assumes your judge can tell right from wrong

The interesting risk in skill-up is not in the runner, it is in the feedback direction. skill-upper reads a failed report and is allowed to repair the eval case as well as the skill. That is the correct call when a case encodes a bad expectation, and it is exactly the wrong call when the case is right and the skill is broken. A loop that can edit its own tests can also edit away its failures.

The README does not describe a guard for that. There is no documented policy for when a case may be changed, no provenance marker separating cases a human wrote from cases the agent rewrote, and no rollback procedure described for an iteration that made things worse. The judge strategies are the practical control you have: rule_based and script judges are deterministic and cheap to reason about, while agent_judge introduces a model into the grading step, which is what you want for subjective output and what you should avoid for anything you can assert mechanically. Pick per case, not per suite.

The second limitation is environmental. Every case runs a real agent engine, so a suite costs engine time and model tokens, and the results depend on the engine, the model and the workspace that skill-up sets up. A suite that passes against Claude Code tells you nothing about Codex until you run it there, and the multi-engine support exists precisely because that comparison is something you have to make yourself. On Windows the README points at a separate Windows guide rather than promising parity, so treat platform behaviour as something to confirm before wiring a suite into a Windows CI runner.

## skill-up against the eval harnesses you already have

The obvious alternative is to keep using whatever your agent vendor ships. Anthropic-style evals.json is a real format, and skill-up imports it through skill-up import or detects it with --auto, which is an admission that the format predates this tool. If your whole workflow lives inside Claude Code and you are happy with the reports that client produces, adding a Go binary between you and the eval run buys you little.

The difference shows up when you want to compare engines or run headless. A vendor harness generally evaluates inside that vendor's client; skill-up treats the engine as a parameter, so the same case set can be pointed at claude_code, codex, qodercli or qwen_code and the reports land in the same shape. The second difference is the report surface: grading.json and benchmark.json for tooling, JUnit XML for CI dashboards, HTML for humans. The third is skill-upper, which has no equivalent in a plain eval format. If none of those three matter to you, the honest answer is that you do not need this tool yet.

One more comparison worth making is against writing your own shell script around an agent CLI. That is genuinely cheaper for a handful of cases, and it is what most teams start with. It stops being cheaper the moment you want per-case judges, machine-readable output and a CI report, which is the point at which the YAML format starts paying for itself.

## Licence, releases and what maintenance actually costs

skill-up is Apache-2.0, which permits commercial use and modification and includes a patent grant. That matters if you intend to vendor the CLI into an internal platform. The licence covers the skill-up code; it does not cover the agent engines you point it at, and those carry their own terms, so a suite that mixes Claude Code and Codex is a suite with two sets of upstream conditions attached. Nothing here is legal advice, and the licence text in LICENSE is the authority.

The repository is not archived, and the last push was on 2026-09-10. Releases have been frequent, with v0.10.0 on 2026-09-01, v0.9.1 on 2026-08-20 and v0.9.0 on 2026-08-12. That cadence cuts both ways for an adopter: fixes arrive quickly, and a pre-1.0 tool can change its YAML schema between minor versions. Pin a version in CI rather than tracking main, and read CHANGELOG.md before upgrading, because your eval.yaml and case files are the artifacts most likely to need edits.

Upgrade cost beyond the schema is mostly the engine matrix. Each supported engine is a separate integration, and the build requires Go 1.25 or later according to the module file. If you build from source, the Makefile exposes build, test, lint and coverage targets and defaults GOPROXY to proxy.golang.org, with a comment noting that mainland China users can override it (GOPROXY=https://goproxy.cn,direct make build). Running the test target uses the race detector, so a full local verification is heavier than a plain go test.

## Conclusion

Adopt skill-up if you already ship a SKILL.md and want pass or fail evidence per behaviour instead of a hand-checked run folder, and if you are willing to run a real agent engine per case. Do not adopt it as a static linter or a docs generator; it does nothing without an engine that can execute the skill. Before trusting a suite, verify three things in your own repository: that your engine's CLI is installed and reachable, that your judge strategy is the right one for the behaviour under test, and that the generated eval.yaml matches the cases you intended, since a repaired case that encodes the wrong behaviour will pass forever.

## FAQ

### What does skill-up do?

It is an evaluation and evolution tool for Agent Skills. The evaluation side runs declarative YAML cases across multiple Agent Engines with rule, script or agent judges and writes structured reports; the evolution side is the skill-upper Agent Skill, which reads failures, repairs or expands the eval suite and reruns skill-up.

### How do I install skill-up?

The README recommends installing skill-upper through the skills package manager first, for example npx skills add https://github.com/alibaba/skill-up/tree/main/skills/skill-upper -g -a codex -y, and notes that skill-upper checks for the CLI and guides the agent through installation. For a manual setup the README gives a shell installer: curl -fsSL https://raw.githubusercontent.com/alibaba/skill-up/main/install.sh | bash.

### Which agent engines does skill-up support?

Qoder CLI, Claude Code and Codex are built-in Agent Engines, and the README also lists qwen_code among the supported engines. User-defined agents are configured through engine.custom, which the README describes as using a local transport and documents in docs/design/custom-engine.md.

### What report formats does skill-up produce?

It outputs Anthropic-compatible grading.json and benchmark.json, plus benchmark.md, result.json, JUnit XML and HTML reports. The JUnit XML output is the one aimed at CI dashboards.

### Is skill-up free to use?

The repository is licensed under Apache-2.0, which permits commercial use and modification and includes a patent grant. The licence covers skill-up itself, not the agent engines you point it at, which have their own terms.

## Sources

- [alibaba/skill-up on GitHub](https://github.com/alibaba/skill-up)
- [License: Apache-2.0](https://github.com/alibaba/skill-up/blob/main/LICENSE)
- [Project website](https://alibaba.github.io/skill-up/)
- [README](https://github.com/alibaba/skill-up/blob/main/README.md)
- [Releases](https://github.com/alibaba/skill-up/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/alibaba-skill-up
