SkillEvaluator: NVIDIA's Three-Tier Framework for Testing Agent Skills
Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.
At a glance
- What is it?
- SkillEvaluator validates, deduplicates and live-tests agent skills in three independent tiers. It is a strong fit for teams shipping skills into a verified catalog, and a heavy one for a single skill in a private repo.
- Who is it for?
- Adopt SkillEvaluator if you publish agent skills to a catalog or marketplace and need repeatable evidence that each one is safe, non-duplicative and actually changes agent behaviour. Do not adopt it for a single private skill you iterate on daily: Tier 3 requires a provider key, an agent CLI credential and a Docker, local or cloud sandbox, and the README calls live evaluation experimental and only for trusted skills and workspaces.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What SkillEvaluator Solves for Skill Authors
An agent skill is a folder of instructions and supporting files that extends an AI agent, per the Agent Skills specification. The hard question is not whether a skill loads. It is whether the skill is well-formed, duplicates guidance the agent already has, and measurably improves agent behaviour. SkillEvaluator is a Python framework that answers those three questions in three tiers, each an independent entry point. Nothing requires running an earlier tier first.
The intended audience is narrow but real: people who maintain skills at a scale where manual review stops working. The project is part of the NVIDIA Verified Skills pipeline, and the tier table reads like a publishing gate rather than a local dev tool. If you maintain one skill for yourself, the framework's structure is more ceremony than you need. If you accept outside contributions into a shared catalog, the split between deterministic checks and model-graded checks is the useful part.
How the Three Tiers Actually Work
Tier 1, validation, answers whether a skill is safe and well-formed. Its deterministic gates cover schema, PII, license, quality, Unicode safety and script linting, and they need no API key. The full security scan is not self-contained: the README states that complete scanner coverage requires the security extra plus external Semgrep, SkillSpector and Gitleaks, and that missing scanner evidence makes validation incomplete. That is an honest admission and also a practical constraint, since those three tools must be installed and kept current separately.
Tier 2 handles deduplication. similarity-check compares a skill against a catalog and needs an embeddings provider. context-optimization-check examines one skill for repeated guidance and needs a chat provider as well. This tier is the one most likely to catch real waste, because overlapping skills consume context that the agent could spend elsewhere.
Tier 3 is live evaluation. According to the README, Harbor powers the sandboxed agent runs, and the tier needs the evaluator provider, the selected agent's credential, and a Docker, local or cloud sandbox. Gating behaviour differs by tier: Tier 1 always gates validate, Tier 2 gates by default unless you pass --no-block-on-dedup, and Tier 3 is advisory unless --block-on-agent-eval promotes its findings into the exit gate. That asymmetry is deliberate and worth understanding before you wire this into CI.
Installing SkillEvaluator and Running a First Validation
The README's quickstart uses uv and installs the evaluation extras from the Git repository. The first command installs the tool against Python 3.13.
uv tool install --python 3.13 "skillevaluator[all] @ git+https://github.com/NVIDIA/SkillEvaluator.git"If your shell cannot find the command afterwards, the README says to run uv tool update-shell and open a new terminal. The second command runs the deterministic gates against any directory containing a SKILL.md. The scoped check list is what keeps this first run keyless.
skillevaluator validate ./my-skill \
--checks schema,pii,license,quality,unicode,lint \
--no-dedupYou should see per-check results for schema, PII, license, quality, Unicode safety and scripts. No API key, Docker daemon or repository clone is required for this run. The project also ships a Dockerfile that installs the public optional dependencies and sets skillevaluator as the entrypoint, so a container build works if you prefer not to touch your host Python.
Provider setup is a single credential for chat and embeddings when you use NVIDIA Build. The README notes that build.nvidia.com offers free inferencing and defaults to nvidia/nemotron-3-nano-30b-a3b.
export SKILL_EVAL_LLM_PROVIDER=nv_build
export NVIDIA_API_KEY='nvapi-...'
skillevaluator models --limit 10The provider options are openai, anthropic, bedrock and openai-compatible, each with its own key variable. One wrinkle: Anthropic and Bedrock do not provide embeddings, so Tier 2 needs a separate OpenAI, NVIDIA Build or OpenAI-compatible embedding provider even when your chat provider is Anthropic.
Where SkillEvaluator Falls Short
The README is unusually direct about Tier 3: it is experimental, and it is only for trusted skills and workspaces, with Docker or cloud recommended for untrusted code. Treat that as a boundary, not marketing hedging. Running arbitrary skill scripts in local mode is the failure mode the warning points at.
Cost is the second constraint. Live model calls and managed sandboxes can incur charges, and the README is precise that local mode avoids managed sandbox charges but not hosted model charges. There is no free path through Tier 3. The project directs readers to a cost planning section before scaling a run, and suggests starting with one agent and a small dataset.
Dependency sprawl is the third. A complete Tier 1 scan depends on three external scanners; Tier 2 depends on an embeddings provider; Tier 3 depends on Harbor, an agent CLI, its credential and a sandbox. That is a lot of moving parts for a tool whose base install is deliberately small. The pyproject comment states the base install must not pull Harbor or LLM clients, which explains the extras layout but also means the interesting checks are never present by default. If your environment cannot run Docker or reach a hosted model endpoint, Tier 3 is simply unavailable to you, and no amount of configuration changes that.
SkillEvaluator Compared with Generic Eval Harnesses
A general agent evaluation framework such as Harbor asks how well a model or agent performs on a task suite. SkillEvaluator asks a narrower question: does this skill change agent behaviour, and is it safe to ship. The two are related by dependency rather than competition, since Harbor powers SkillEvaluator's sandboxed Tier 3 runs.
The practical difference shows up in the artifact under test. A generic harness treats the agent as the subject and the task as fixed. SkillEvaluator treats the skill folder as the subject and generates the tasks around it. create-eval-dataset builds a four-bucket dataset for a skill, and autopilot will create one initial case at evals/evals.json if no accepted evaluation source exists, reusing the file if it already does. That is a materially different workflow from pointing an eval harness at a benchmark.
If you already run a homegrown eval suite, the overlap is Tier 3, and the reason to switch is the deterministic tiers above it. If you only need pass rates on a fixed task set, SkillEvaluator's validation and deduplication machinery is dead weight.
Maintenance, Licensing and Upgrade Cost
The repository is not archived. Its last push was on 2026-09-10, ten days before this writing, and the most recent release is v0.1.0 from 2026-08-05. The pyproject.toml in the repository already declares version 0.2.1, so the source tree is ahead of the published release. That gap matters if you pin versions: installing from the Git URL gets you code that no release tag describes.
Version pinning is uneven. The base dependencies use bounded ranges such as click>=8.3.3,<9, and litellm is capped below 1.89.0.dev0, but several entries like jinja2>=3.1 and pydantic>=2.11.0 have no upper bound. A fresh install months from now may resolve differently than today's. The uv.lock file in the repository is the reproducible path; the README's uv tool install command does not reference it.
Licensing is Apache-2.0, declared in pyproject.toml with license-files pointing at LICENSE, NOTICE and THIRD_PARTY_NOTICES.md. The practical implication is that you inherit the third-party notice obligations of the bundled dependencies, and Tier 1 can check the license of a skill you evaluate but not of SkillEvaluator itself. This is not legal advice; review THIRD_PARTY_NOTICES.md if you redistribute the tool inside a product.
Running the Full Pipeline Before You Commit
Once Tier 1 passes, the README's path to a complete run is doctor followed by validate --full. Both commands are shown with codex as the agent and docker as the environment mode.
skillevaluator doctor --agents codex --env-mode docker
skillevaluator validate ./my-skill \
--full \
--agents codex \
--env-mode dockerThe doctor command verifies the selected agent runtime before you spend anything on live runs. The --full flag runs all three tiers and enables autopilot. For a wider dataset than autopilot's single initial case, the README suggests generating and reviewing the four-bucket set first.
skillevaluator create-eval-dataset ./my-skill --fullTwo things to watch. Autopilot writes to evals/evals.json and reuses that file on later runs, so an unreviewed first case becomes your baseline unless you inspect it. And if you want Tier 2 findings to stop failing the pipeline, --no-block-on-dedup keeps the scan and reports but makes the results advisory.
Editorial conclusion
Adopt SkillEvaluator if you publish agent skills to a catalog or marketplace and need repeatable evidence that each one is safe, non-duplicative and actually changes agent behaviour. Do not adopt it for a single private skill you iterate on daily: Tier 3 requires a provider key, an agent CLI credential and a Docker, local or cloud sandbox, and the README calls live evaluation experimental and only for trusted skills and workspaces. Before committing, run skillevaluator doctor --agents codex --env-mode docker against your chosen runtime, confirm Semgrep, SkillSpector and Gitleaks are installed so scanner evidence is complete, and check the Tier 3 guide's cost section, since hosted model calls and managed sandboxes can incur charges.
Frequently asked questions
What does skill evaluation mean in SkillEvaluator?
SkillEvaluator treats evaluation as three separate questions: whether a skill is safe and well-formed, whether it overlaps with skills already in a catalog, and whether it changes live agent behaviour. The tiers are independent entry points, so you can run only the one you need.
How do you evaluate an agent skill with SkillEvaluator?
Start with skillevaluator validate against a directory containing SKILL.md, scoping the checks to schema, pii, license, quality, unicode and lint for a keyless first run. Deeper tiers add similarity-check, context-optimization-check and validate --full, which require provider credentials and, for Tier 3, an agent CLI and a sandbox.
Does SkillEvaluator need an API key?
Not for the deterministic Tier 1 checks, which the README's quickstart runs without any API key, Docker daemon or repository clone. Tier 2 needs an embeddings provider, and Tier 3 needs the evaluator provider plus the selected agent's credential.
Which Python versions does SkillEvaluator support?
The pyproject.toml declares requires-python as >=3.12,<3.14, and the classifiers list Python 3.12 and 3.13. The README's quickstart installs with --python 3.13.
Is SkillEvaluator free to run?
The tool is Apache-2.0 licensed, and Tier 1 deterministic checks cost nothing beyond the external scanners. The README warns that live model calls and managed sandboxes can incur charges, and that local mode avoids managed sandbox charges but not hosted model charges.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/nvidia-skillevaluator)