NVIDIA SkillSpector: a static and LLM scanner for agent skills you are about to install
Security scanner for AI agent skills. Detect vulnerabilities, malicious patterns, and security risks.
At a glance
- What is it?
- SkillSpector answers one question before you add a skill to Claude Code, Codex CLI or Gemini CLI: is this safe to install? It runs 71 pattern checks, optional LLM semantic analysis, and ships as a Python CLI, a Docker image and a Pi extension.
- Who is it for?
- Adopt SkillSpector if you install third-party agent skills and want a gate before they run with implicit trust, and especially if you already emit SARIF from CI. Skip it if you need a stable API: pyproject.toml still classifies the project as Development Status :: 3 - Alpha, and the suppression and baseline formats are the parts most likely to shift between minor versions.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
The trust gap SkillSpector is built to close
Agent skills are markdown and code that an agent loads and then executes with the same trust it gives its own instructions. The README states the premise directly: skills "execute with implicit trust and minimal vetting", and it cites research figures of 26.1% of skills containing vulnerabilities and 5.2% showing likely malicious intent. You do not have to accept those numbers to see the structural problem. A skill is a dependency, but it rarely goes through the review a package manager dependency gets, because it often arrives as a folder someone shared.
SkillSpector is aimed at the person deciding whether to add that folder. It is a Python command line tool that takes a skill from a Git repository, a URL, a zip file, a directory or a single SKILL.md file and produces a risk score from 0 to 100 with severity labels and recommendations. It is also one stage of the NVIDIA Verified Skills pipeline, which scans, evaluates and signs skills before publication to the NVIDIA skills catalog. That second role matters for how you read the tool: it is designed to run unattended as a gate, not only as an interactive aid. The README frames the goal as answering "Is this skill safe to install?"
Two analysis stages and where the checks come from
The scanner runs in two stages. The first is static analysis over the ingested files, and it is the stage that always runs. The README lists 71 vulnerability patterns across 17 categories: prompt injection, data exfiltration, privilege escalation, supply chain, excessive agency, output handling, system prompt leakage, memory poisoning, tool misuse, rogue agent, anti-refusal, trigger abuse, dangerous code found through AST inspection, taint tracking, YARA signatures, MCP least privilege and MCP tool poisoning. The spread is the interesting part. Several of these categories describe things a conventional SAST tool has no vocabulary for, because the artifact is a set of instructions to a model rather than a program. Anti-refusal and trigger abuse, for instance, are about how a skill steers the agent, not about a memory bug.
The second stage is optional LLM semantic evaluation, and it is what the --no-llm flag turns off. This is the design decision worth pausing on. Static patterns are cheap, deterministic and explainable, but they miss intent expressed in prose. An LLM pass can read the skill and judge it, at the cost of sending the skill's contents to a provider and paying per scan. The .env.example shows the provider list is long: openai, anthropic, anthropic_proxy, bedrock, nv_build, ollama, azure_openai, openai_compatible, claude_cli, codex_cli and gemini_cli, with nv_build as the default. Ollama on that list means a fully local LLM stage is a supported configuration, which is the version of this tool that a security team with strict data rules can actually run.
A third piece sits alongside the two stages: pattern SC4 queries OSV.dev for live CVE data, with an automatic offline fallback. That is a network dependency with a defined degraded mode rather than a hard requirement, which is the right shape for a scanner that may run in an air-gapped build.
Installing SkillSpector and running a first scan
The README offers three routes. The fastest is uv, which installs the CLI as a tool. If you intend to use the MCP server, the MCP extra has to be present at install time, so use the second form instead of adding it later.
uv tool install git+https://github.com/NVIDIA/skillspector.git
uv tool install 'skillspector[mcp] @ git+https://github.com/NVIDIA/skillspector.git'From source, the Makefile targets assume an active virtual environment, and the Makefile uses uv when it is available and pip otherwise.
git clone https://github.com/NVIDIA/skillspector.git
cd skillspector
uv venv .venv && source .venv/bin/activate
make installIf you would rather not install Python at all, the repository includes a Dockerfile based on the Docker Official Python 3.12-slim-bookworm image. The container's working directory is /scan, so you mount your current directory there and the path you pass to the scanner is relative to that mount.
make docker-build
docker run --rm -v "$PWD:/scan" skillspector scan ./my-skill/ --no-llmFor a first real scan, start without the LLM stage so you see what the static patterns alone produce. The default output is a formatted terminal report; --format json, markdown or sarif writes to a file instead.
skillspector scan ./my-skill/ --no-llm
skillspector scan ./my-skill/ --no-llm --format json --output report.jsonWhat you should see is a report with a risk score, severity labels and recommendations. To add the semantic stage, set SKILLSPECTOR_PROVIDER and the matching credential, or leave the provider unset and set NVIDIA_INFERENCE_KEY for the default nv_build path, then drop the --no-llm flag. Two operational caps are worth knowing before you point this at a large repository: SKILLSPECTOR_MAX_WORKFLOW_SECONDS defaults to 600 for a whole scan workflow, and SKILLSPECTOR_MAX_STATIC_ANALYSIS_SECONDS_PER_ARTIFACT defaults to 300 per artifact, capped by whatever remains of the workflow budget.
Ingest limits, fail-closed behaviour and the cost of scanning untrusted input
A scanner that downloads arbitrary URLs and unzips arbitrary archives is itself an attack surface, and the README treats it that way. Two independent caps apply to remote and archive inputs: INGEST_MAX_BYTES at 100 MiB, covering streamed URL downloads, the total uncompressed size of a zip and post-clone disk usage of a Git repository, and INGEST_MAX_ZIP_MEMBERS at 10,000, capping entries in a single zip. Breaching either raises IngestLimitExceededError, which the README describes as failing closed.
The README is careful to distinguish these from a third limit, MAX_FILE_BYTES at 1 MB per file, which is downstream: it bounds what individual analyzers read out of a directory that has already been ingested. The distinction is easy to miss and it matters when you are tuning. Raising the per-file cap does not raise the ingest ceiling, and hitting the ingest ceiling is an error rather than a truncated scan. If your scan of a monorepo-style skill fails with an ingest error, the fix is to scan a narrower path, not to adjust MAX_FILE_BYTES.
There is a further set of ceilings documented separately in docs/ANALYSIS_RESOURCE_BOUNDS.md, covering bundles, parsers, nested artifacts, the ledger and findings. The README does not reproduce those numbers, so the honest position is that the analysis-side bounds live in that file and should be read before you rely on a scan completing in a fixed time on adversarial input.
Baselines, batch scanning and CI integration
A scanner that reports the same 40 findings on every run gets ignored, so suppression is the feature that decides whether SkillSpector survives in a real repository. The README describes accepting known findings through either a glob-rule or a fingerprint baseline, so that re-scans surface only new issues, with details in docs/SUPPRESSION.md. The repository ships .skillspector-baseline.example.yaml at the top level as the starting shape. The trade-off is inherent: a baseline is a promise that a finding has been reviewed, and once the file exists, nothing in the tool stops it from absorbing a genuinely new problem that happens to match an old rule.
For many skills at once, contrib/batch_scan/ runs scans in parallel. The README gives three invocations, including a --workers 20 example, and notes multilingual detection for zh, ja and ko. Higher LLM concurrency is handled by configuring multiple API keys following contrib/batch_scan/.env.example. This is the part of the project that looks least like a finished product: it lives under contrib/ rather than in the installed package, and you invoke it as a module rather than through the skillspector entry point, which is worth knowing before you build a pipeline on it.
SARIF output is the CI story. The README lists SARIF specifically for CI/CD integration and IDE tooling, and the Dockerfile's ENTRYPOINT is skillspector, so the container can be dropped into a job as-is. The MCP extra and the Pi extension in extensions/skillspector.ts cover the other direction: scanning from inside an agent session rather than before it.
Where SkillSpector is the wrong tool
The clearest limitation is that the LLM stage is optional and the static stage carries the load. If you run with --no-llm, you get pattern matching, AST checks, taint tracking and YARA signatures, and nothing that reasons about what a skill is trying to do in prose. A skill that reads innocuously and steers an agent toward a bad outcome is exactly the class the semantic stage exists for, and it is the class the static stage is weakest on. Turning the LLM stage on has its own cost: skill contents leave your machine unless you point SKILLSPECTOR_PROVIDER at ollama or another local endpoint.
Second, the project is self-described as alpha. pyproject.toml carries Development Status :: 3 - Alpha, and the release history shows three releases inside eleven days in August 2026, with the last push on 2026-08-28. Rapid iteration on a security tool is not automatically bad, but it means rule IDs, suppression formats and report schemas are moving. Anything you build that parses the JSON output should pin the version.
Third, this is a pre-installation scanner, not a runtime monitor. It tells you what a skill contains at scan time. It does not observe what the skill does once the agent is running, and a skill that fetches instructions at runtime can look clean at scan time. If your threat model includes a skill that changes behaviour after installation, SkillSpector is one input, not the answer. Finally, the dependency list is not small: langgraph, langchain-anthropic, langchain-aws, langchain-openai, boto3, langsmith, openai and yara-python all sit in the core dependencies, alongside a pinned typer range that the pyproject comments explain is constrained by a click version conflict with semgrep. That is a lot of surface for a tool you are running against untrusted input.
How SkillSpector differs from Semgrep and generic SAST
The natural comparison is Semgrep or another general-purpose static analyzer pointed at a skill directory. The difference is in what each tool models. Semgrep matches syntactic and semantic patterns in source code across many languages, and its rule ecosystem is built around code constructs. SkillSpector's categories include prompt injection, system prompt leakage, anti-refusal, trigger abuse and MCP tool poisoning, which are properties of instructions and tool declarations rather than of executable code. A Semgrep rule can flag a suspicious string, but it has no notion of a skill's declared capabilities or of an agent's refusal behaviour.
The second difference is the pipeline around the scan. SkillSpector produces a 0-100 risk score with severity labels and recommendations, queries OSV.dev for CVE data, and supports a baseline file for suppressing known findings. Semgrep gives you findings and lets you build the rest. If you already run Semgrep in CI, the overlap is real but partial: Semgrep covers the dangerous-code category better and more broadly, and SkillSpector covers the agent-specific categories that Semgrep has no rules for. Running both is defensible; running only Semgrep on agent skills leaves the instruction-layer categories unexamined.
Licence, upgrades and what maintenance looks like
SkillSpector is Apache-2.0, and pyproject.toml declares license = "Apache-2.0" with the matching OSI classifier. The README adds an open-source software notice stating that the project will download and install additional third-party open source projects and that you should review their licence terms before use. That notice is not boilerplate here: the dependency list pulls in LangChain packages, boto3, langsmith and yara-python, and the Docker image installs git and ca-certificates on top of python:3.12-slim-bookworm. Apache-2.0 covers SkillSpector itself; it does not cover what those dependencies bring, and THIRD_PARTY_NOTICES.md at the repository root is where that inventory lives. Nothing here is legal advice, and a team shipping a commercial product should read that file rather than assume the permissive licence of the top-level project settles the question.
Upgrade cost is tied to how you installed it. A uv tool install is updated with uv tool update skillspector, which is a single command but also an unpinned jump to whatever is current. That is a poor fit for a tool whose rule set and output schema are changing between minor versions. Python support is declared as >=3.12,<3.15 with classifiers for 3.12, 3.13 and 3.14, so the interpreter ceiling is explicit and worth checking before a platform upgrade. The repository includes a .pre-commit-config.yaml, a CHANGELOG.md and a CONTRIBUTING.md, which is the minimum you want to see before depending on a scanner in CI. The README does not document a rollback procedure or a long-term support policy, so pinning is your mechanism, not the project's.
Editorial conclusion
Adopt SkillSpector if you install third-party agent skills and want a gate before they run with implicit trust, and especially if you already emit SARIF from CI. Skip it if you need a stable API: pyproject.toml still classifies the project as Development Status :: 3 - Alpha, and the suppression and baseline formats are the parts most likely to shift between minor versions. Before rolling it into a pipeline, run one scan with --no-llm on a skill you already trust, then a second with your provider key set, and compare the two reports; the LLM stage is where findings and cost diverge. Pin the version with uv tool install and re-read CHANGELOG.md before each bump.
Frequently asked questions
Who are the main AI agents that SkillSpector targets?
The README names Claude Code, Codex CLI and Gemini CLI as the agents whose skills it scans, and the pyproject description adds Cursor and similar tools. It scans the skill artifact itself, so the same check applies whichever of these agents loads it.
What is the main purpose of an AI agent, according to SkillSpector's documentation?
The README does not define the purpose of an AI agent. Its focus is the skills an agent loads, which it says execute with implicit trust and minimal vetting, and the scanner exists to answer whether a given skill is safe to install.
Which tasks is SkillSpector most suitable for?
It is built for checking an agent skill before installation, across Git repositories, URLs, zip files, directories or a single SKILL.md file. It also serves as one stage of the NVIDIA Verified Skills pipeline, which scans, evaluates and signs skills before publication to the NVIDIA skills catalog.
What are the five parts of an AI agent?
The README does not describe a five-part model of an agent. What it does enumerate is the scan pipeline: two stages, static analysis and optional LLM semantic evaluation, plus an OSV.dev lookup for live CVE data with an offline fallback.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/nvidia-skillspector)
Community notes