Model or dataset
darkrishabh/agent-skills-eval avatar
darkrishabh/agent-skills-eval

Without the baseline flag, agent-skills-eval measures nothing

A test runner for agentskills.io-style AI agent skills

797 stars43 forksTypeScriptMIT

At a glance

What is it?
agent-skills-eval runs every prompt twice against the same model, once with your SKILL.md in context and once without it, then lets a judge model grade both sides independently. That is the whole idea, and it is off by default, which is the first thing to know before you read a report it produces.
Who is it for?
This is a good fit if you ship SKILL.md files and have never been able to show whether they help, because the with_skill against without_skill comparison plus JSON artifacts is exactly the missing evidence. It is not a fit if you want an offline rubric check with no API calls, since grading goes through a chat model every time, and the cost of a run is the target call times the number of evals times two.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 5 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The comparison is the product, and it ships switched off

The mechanism is simple enough to state in one line: for every eval defined in your skill, the same prompt is sent twice, once with `SKILL.md` loaded into context and once as a bare baseline, and a judge model grades each side separately against the eval's `expected_output` and `assertions`. The judge never sees both answers together, which is what keeps the grading independent rather than comparative. What matters is the default: `baseline` is a boolean whose default is `false`, and the page is explicit that without it you only get the `with_skill` run. A report from a default invocation shows you outputs, timings and grades, and no lift at all.

bash
npx agent-skills-eval ./skills \
  --target gpt-4o-mini \
  --judge gpt-4o-mini \
  --baseline \
  --strict

The default judge is the same model as the model under test

Both `target` and `judge` default to `gpt-4o-mini`, and the options table addresses this directly: set the judge to a stronger model than the target for more reliable grading. Left alone, the out-of-the-box run has a model grading its own output, which is the configuration least able to notice that a skill made the answer worse. Two other defaults push in the same cautious direction. `targetParams` and `judgeParams` both set `temperature: 0`, and they are the two options the page marks as config-only, so you cannot pass them as flags and have to change them in the file. Zero temperature removes sampling noise from both sides. It does nothing about a judge that is agreeable by construction.

baseUrl has no fallback, and apiKeyEnv names a variable instead of a key

The credential handling is the cleanest part of the configuration. `baseUrl` defaults to the `OPENAI_BASE_URL` environment variable and is otherwise required, so there is no baked-in endpoint to be surprised by. `apiKeyEnv` defaults to `OPENAI_API_KEY` and holds the name of an environment variable rather than the key itself, and the page states plainly that the key is never written to config. That is what lets the same YAML file run against OpenAI, Together, Groq, Anthropic through its OpenAI-compatible layer, or a local Llama server, since everything is addressed through the OpenAI chat API shape.

bash
OPENAI_API_KEY=... npx agent-skills-eval --config agent-skills-eval.yaml

CLI flags always override config values, so a CI job can point the same file at a different endpoint without editing it.

The package exports an agent runtime the page says it does not need

The positioning paragraph says the tool is the test framework for the Agent Skills ecosystem, separated from any specific agent runtime so it works wherever your skills do. The export map says something more complicated. Alongside the root entry, `./provider`, `./openai-compatible`, `./config` and `./reporters`, there is an `./experimental/runtime` subpath that resolves to a bundled `agent-runtime.js` with its own type declarations. So a runtime adapter ships in the tarball, labelled experimental, which is a different claim from having no runtime at all. The published `files` list is equally selective: it ships `docs/artifact-contract.md` and `docs/runtime-foundation.md` by name, while the rest of the `docs/` tree in the repository is left behind. Two of the documents that define the output contract are therefore available only inside the installed package.

include and exclude are globs, and exclude is applied second

Skill discovery starts at `root`, whose default is the current directory and which is scanned recursively for `SKILL.md` files; on the command line it is the positional argument, as in the quickstart. From there, `include` defaults to all discovered skills and `exclude` defaults to none, both are globs matched against each skill's path, and exclude is applied after include, so a skill matched by both is skipped. Both are repeatable as `--include` and `--exclude`. Inside a selected skill, `evalIds` picks which cases run, defaulting to all of them, and the behaviour on a miss is deliberate: missing IDs fail the run rather than being silently skipped. The SDK also accepts numeric IDs where the CLI takes strings.

Artifacts are JSON and JSONL all the way down to a static page

A run writes a workspace laid out as iteration directories, so the second run sits beside the first instead of overwriting it:

text
agent-skills-workspace/
└── iteration-1/
    ├── meta.json            # run metadata
    ├── benchmark.json       # rolled-up pass/fail per skill
    ├── eval-basic/
    │   ├── with_skill/      # output, timing, judge grading
    │   └── without_skill/   # ↑ same, with the skill stripped
    └── report/
        └── index.html       # the visual report

The `layout` option is what produces that `iteration` naming. Logging has three formats, `pretty`, `jsonl` and `silent`, driven by `--log-format`, `--log-file`, `--verbose` and `--no-color`, and the report is switched with `--report`, `--no-report`, `--report-title` and `--report-output`. Everything else is `grading.json`, `benchmark.json` and `meta.json` per eval, so a dashboard can read the JSON and ignore the HTML entirely.

Tests run against the build, and the tarball ships a sample skill

The test script is `npm run build && node --test test/*.test.mjs`, which means the suite exercises the compiled output in `dist` rather than the TypeScript sources, and a broken build fails the tests before any test runs. `prepublishOnly` runs `npm test`, so that gate sits in front of a publish, and `prepack` rebuilds first. The published `files` array is worth reading for a different reason: it includes the whole `examples/` directory, so `examples/agent-skills-eval.yaml` and `examples/basic-skill/` arrive inside the npm package. The version is 0.1.1, the repository has no GitHub releases, and the most recent push was 2026-09-30, so the changelog file rather than the release list is where the change history lives.

Editorial conclusion

This is a good fit if you ship SKILL.md files and have never been able to show whether they help, because the with_skill against without_skill comparison plus JSON artifacts is exactly the missing evidence. It is not a fit if you want an offline rubric check with no API calls, since grading goes through a chat model every time, and the cost of a run is the target call times the number of evals times two. Before trusting a report, turn on the baseline, point the judge at a stronger model than the target, and keep both temperatures at the configured zero so a difference in the numbers is a difference in the skill. Read `docs/artifact-contract.md` from the package before wiring the output into a dashboard.

Frequently asked questions

What is agent-skills-eval and what does it measure?

It is a TypeScript SDK and CLI for evaluating Agent Skills, the open standard from Anthropic for giving agents domain knowledge. It runs each eval twice, once with the SKILL.md in context and once as a baseline, and has a judge model grade both sides so you can see whether a skill changes the output rather than assuming it does.

How do I install agent-skills-eval?

Install it with `npm install agent-skills-eval`, or run it without installing via `npx agent-skills-eval --help`. The binary is `dist/cli.js` and the package exposes its SDK from the root entry, so the same install gives you both the command line runner and the library.

Which providers can agent-skills-eval use?

Anything that speaks the OpenAI chat API, which the page names as OpenAI, Together, Groq, Anthropic through OpenAI-compatible layers, and local Llama servers. The `baseUrl` option has no built-in endpoint, defaulting to the `OPENAI_BASE_URL` environment variable and otherwise being required.

Why does my agent-skills-eval report show no improvement?

The most likely cause is that `baseline` is false, which is its default, and in that state only the `with_skill` run happens and there is nothing to compare against. Pass `--baseline` or set `baseline: true` in the config, and check that `judge` is set to a stronger model than `target` rather than defaulting to the same gpt-4o-mini.

Does agent-skills-eval follow the agentskills.io specification?

The page states it implements the full specification, including SKILL.md validation, the `evals/evals.json` layout, the official `iteration-N` artifact layout and the frontmatter rules. It is described as a test framework for that ecosystem rather than a runtime for a specific agent.

Official sources

  1. darkrishabh/agent-skills-eval on GitHub
  2. Issues
  3. License: MIT
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/darkrishabh-agent-skills-eval.svg)](https://hysenlabs.com/projects/darkrishabh-agent-skills-eval)