Model or dataset
mgechev/skillgrade avatar
mgechev/skillgrade

Skillgrade: unit tests for the agent skills you wrote

"Unit tests" for your agent skills

721 stars47 forksTypeScriptMIT

At a glance

What is it?
Skillgrade runs an agent against your SKILL.md in a container, scores each trial with deterministic scripts or LLM rubrics, and reports a pass rate. It is for skill authors who want evidence that an agent discovers and uses their skill, not just that the file exists.
Who is it for?
Adopt Skillgrade if you maintain a SKILL.md and want a number attached to it: the eval.yaml format, the deterministic grader channel through SKILLGRADE_INPUT, and the --ci threshold all exist to make skill quality checkable in a pipeline. Do not adopt it if you have no Docker available, or if you want to evaluate a full agent product rather than one skill, since the model is one skill, one task, one score.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 34 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

The problem Skillgrade solves: a skill file is not evidence

A SKILL.md is a promise. It tells an agent when to reach for a capability and how to use it. Nothing about writing that file proves an agent will notice it, pick it over the tools it already has, or follow the steps in the order you wrote them. Skillgrade's README frames the tool as tests that check AI agents correctly discover and use your skills, which is a narrower and more useful claim than evaluating an agent in general. The unit under test is the skill, not the model. That distinction decides everything else about the tool: the scaffolding lives in your skill directory, the tasks describe work your skill is supposed to enable, and the score is a pass rate across repeated trials rather than a single transcript someone reads and nods at. The intended user is the person who publishes a skill and gets asked whether it works. Without a harness, the honest answer is that it worked when I tried it. With one, the answer is a percentage over N trials at a stated agent and model.

How a run is structured: eval.yaml, containers, trials and graders

An eval.yaml declares defaults and a list of tasks. Defaults cover the agent, the model, the provider (docker or local), the trial count, a timeout in seconds, a CI threshold, and the default grader model and provider. Each task carries a name, an instruction, an optional workspace mapping of source files to destinations inside the container, and a list of graders with weights.

Two grader types exist. A deterministic grader runs a command such as npx ts-node graders/check.ts and returns a score. An llm_rubric grader takes a rubric string and asks a model to judge the transcript. Weights let one task combine both, for example a script at 0.7 and a rubric at 0.3. The README also documents per-task overrides for agent, model, grader_provider, trials and timeout, so one hard task can run ten trials against Claude while the rest of the suite stays cheap.

The data flow around scoring is the part worth reading closely. A task may declare expected, described as the answer key, and metadata labels for filtering. Skillgrade delivers expected and never interprets it; comparison is the grader's job. Both are handed to graders only after the agent process has exited, and neither is written into the workspace. A deterministic grader receives one JSON document in the environment variable SKILLGRADE_INPUT containing task, trial, expected and metadata. An llm_rubric grader instead sees expected rendered as an ## Expected Output section in its prompt. That split is deliberate: the agent under test cannot read the answer key, and a grader script does not have to know the shape of every task, since the same script can compare any token task by reading expected.token.

Installing Skillgrade and running a first smoke test

Prerequisites are Node.js 20 or newer and Docker. The package installs globally from npm:

bash
npm i -g skillgrade

Change into the skill directory, which must contain a SKILL.md, and scaffold an eval.yaml. The API key you pass selects the agent, so GEMINI_API_KEY targets Gemini, ANTHROPIC_API_KEY targets Claude, and OPENAI_API_KEY targets Codex.

bash
cd my-skill/
GEMINI_API_KEY=your-key skillgrade init

According to the README, init generates eval.yaml with AI-powered tasks and graders. Without an API key it writes a well-commented template instead. The --force flag overwrites an existing eval.yaml, so use it only when you mean to discard your edits.

After customizing eval.yaml, run the smoke preset, which is 5 trials:

bash
GEMINI_API_KEY=your-key skillgrade --smoke

Then read the result. The CLI report is the default preview, and the browser view serves a web UI on port 3847:

bash
skillgrade preview
skillgrade preview browser

Reports land in $TMPDIR/skillgrade/<skill-name>/results/, and --output=DIR moves them elsewhere. The other presets are --reliable at 15 trials and --regression at 30, which the README describes as a high-confidence regression detection setting. Expect a smoke run to be fast and noisy: 5 trials is a capability check, not an estimate you should quote.

Writing a grader that reads SKILLGRADE_INPUT

The README's token-check example is the clearest statement of the grader contract. The script parses the environment variable, reads an answer file from the workspace, and prints a JSON object with score and details on stdout.

js
// graders/check-token.mjs
const { task, trial, expected, metadata } = JSON.parse(process.env.SKILLGRADE_INPUT);
const answer = JSON.parse(fs.readFileSync('answer.json', 'utf8')); // cwd is the workspace
console.log(JSON.stringify({
  score: answer.token === expected.token ? 1 : 0,
  details: `${task}: want ${expected.token}, got ${answer.token ?? '(none)'}`,
}));

Two details matter more than they look. The working directory during grading is the workspace, so relative paths in a grader point at files the agent produced. And the answer key stays out of that workspace, which means a grader can be strict without leaking the solution into the agent's context. Shell graders get the same document, and the README shows jq reading it directly:

bash
want=$(jq -r .expected.token <<< "$SKILLGRADE_INPUT")

If you have written graders before, the discipline here is familiar: emit a numeric score, keep the comparison in your script, and put anything a human needs to debug into details. Tasks without expected are still valid and are scored purely on what their graders measure, so metric-style tasks and golden-truth tasks can share one suite.

Where Skillgrade is the wrong tool

The trial model is the main constraint. Every run starts a fresh agent session against a container, so a suite of 30 trials across several tasks multiplies into real API spend and real wall-clock time. The README's presets exist precisely because 5 trials is cheap and 30 is not. If your skill is expensive to exercise, or if your tasks depend on a service the container cannot reach, the cost of a trustworthy pass rate may exceed what the skill is worth.

Docker is a hard prerequisite in the documented setup, and the docker provider is the default. The --provider=local option exists, but the container is where the workspace mapping, the chmod on copied binaries, and the resource limits (cpus, memory_mb) actually take effect. Running locally changes what you are testing, and the README does not document rollback or cleanup semantics for either provider, so treat state left behind by a run as something to check yourself.

The scoring model is also narrower than it first appears. Skillgrade delivers expected and never interprets it. That is a clean separation, but it means a wrong pass rate is almost always a grader bug or a task that does not isolate the behaviour you care about, not something the tool can detect for you. There is no documented mechanism for detecting a grader that always returns 1. The --validate flag, which the options table describes as verifying graders using reference solutions, is the closest thing to a guard against that, and it is worth understanding before you trust a number in CI.

How Skillgrade differs from Langfuse and SkillsBench

Langfuse and SkillsBench appear in the searches around this project, and the comparison is instructive even though the README does not discuss either. Langfuse is an observability platform: you instrument an application, traces flow in, and you inspect and score production traffic after the fact. Skillgrade runs before deployment, in a container, against a fixed task list, and its output is a pass rate rather than a trace stream. If your question is why did this session go wrong, an observability tool is the right shape. If your question is does this skill work at all, and did my last edit break it, you want a suite that runs the same tasks repeatedly and compares numbers.

SkillsBench-style evaluation sits at the other end: a fixed benchmark of tasks and skills, useful for comparing models or approaches on common ground. Skillgrade is the opposite: your skill, your tasks, your graders, your threshold. The trade-off is that nothing you produce is comparable to anyone else's number. That is fine for a regression gate in your own repository and useless for a leaderboard. The --ci flag with --threshold=0.8 is the tell: the tool is built to fail a build, not to publish a score.

Maintenance, licence and what you are taking on

The repository is not archived, and the last push was on 2026-08-26. The package is version 0.3.0, which is pre-1.0, and the eval.yaml format is versioned separately with version: "1" at the top of the file. That version field is your signal that the schema can move; a scaffolded file records the format it was written against, so regenerating with skillgrade init --force after an upgrade is a reasonable habit, provided you keep your own copy to diff against.

The licence is MIT, stated in both the repository and package.json. MIT is permissive, so you can use Skillgrade in a commercial repository and vendor it if you need to. That is a statement about the licence text, not advice about your situation; the obligations you actually carry depend on how you distribute anything derived from it.

The upgrade cost is mostly in the graders, not the CLI. Deterministic graders are ordinary scripts reading SKILLGRADE_INPUT, so they survive CLI upgrades as long as that contract holds. LLM rubrics are the fragile part: they depend on a grader model and provider, and the README's own examples use preview model names such as gemini-3-flash-preview and gemini-3.5-flash. Pinning grader_model in defaults is the difference between a rubric that scores consistently and one that quietly drifts when a provider rotates a model.

Editorial conclusion

Adopt Skillgrade if you maintain a SKILL.md and want a number attached to it: the eval.yaml format, the deterministic grader channel through SKILLGRADE_INPUT, and the --ci threshold all exist to make skill quality checkable in a pipeline. Do not adopt it if you have no Docker available, or if you want to evaluate a full agent product rather than one skill, since the model is one skill, one task, one score. Verify first that your skill has a SKILL.md at its root, that a --smoke run of 5 trials produces a pass rate you can explain, and that your graders read expected from SKILLGRADE_INPUT rather than parsing stdout.

Frequently asked questions

What is a skill description in Skillgrade?

The README does not define a separate skill description field; it treats the skill as a directory containing a SKILL.md, which skillgrade init looks for when scaffolding eval.yaml. The description of what a task is asking for lives in that task's instruction field.

What is the skill of evaluation in Skillgrade?

Skillgrade separates evaluation into two grader types: deterministic graders that run a command and return a score, and llm_rubric graders that judge the transcript against a rubric. Weights on each grader decide how much of the task score comes from each.

What is a rating scale for skills in Skillgrade?

There is no fixed rating scale. Each grader returns a score, weights combine them per task, and the suite reports a pass rate; the --threshold value, 0.8 in the README's example, is what --ci compares that pass rate against.

What does skill level mean in Skillgrade?

Skillgrade does not assign skill levels. It reports a pass rate over repeated trials, with presets of 5 trials for --smoke, 15 for --reliable and 30 for --regression, and the README warns that the trial count determines how much confidence the number deserves.

Official sources

  1. Issues
  2. License: MIT
  3. mgechev/skillgrade on GitHub
  4. Project website
  5. README
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/mgechev-skillgrade.svg)](https://hysenlabs.com/projects/mgechev-skillgrade)
Community notes

Community notes