cisco-ai-defense/skill-scanner: a best-effort scanner for AI Agent Skills
Security Scanner for Agent Skills
At a glance
- What is it?
- Skill Scanner combines YAML and YARA-X signatures, AST and dataflow analysis, an optional LLM judge and a CEL decision layer to flag prompt injection and exfiltration in Codex and Cursor skills. Its own evaluation says the detection gate is not yet passed.
- Who is it for?
- Adopt Skill Scanner if you ship or consume Codex and Cursor skills and want signature, dataflow and optional LLM checks wired into CI through SARIF output and exit codes. Do not adopt it as a certification gate: the README states that no findings does not mean no risk, and the locked source-disjoint split reports 7.75% recall with a 7.71% false-positive rate, so bundled CEL rules stay in shadow.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 4 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem: agent skills are instructions, and instructions can be hostile
An Agent Skill is mostly text. OpenAI Codex Skills and Cursor Agent Skills follow the Agent Skills specification, and what they carry is a description plus instructions the agent will follow at runtime. That is a different threat model from a library that runs in your process. A skill can tell the model to read a file and send it somewhere, or to treat attacker-controlled content as a command. There is no compiler and no type system standing between the skill author and the agent's behaviour.
Skill Scanner exists for that gap. It is aimed at people who publish skills, and at teams that pull skills from a registry or a shared repository into their own agent setup. The README frames it as best-effort detection, not coverage: a clean scan means no known pattern matched, and the project repeats that warning in its own limitations section. That framing matters more than the feature list. A scanner for natural-language instructions cannot be complete, and this one does not claim to be.
How the detection pipeline is assembled
The scanner layers several engines rather than relying on one. Pattern-based detection uses YAML rules and YARA-X. On top of that sit AST and dataflow analysis, which look at what a skill's code actually does with data rather than only at strings. An optional LLM-as-a-judge adds semantic analysis for cases where a signature would not fire. A bounded CEL decision layer sits between deterministic detection and the optional LLM stage, correlating typed detector facts.
The CEL piece is the most interesting design choice and the most hedged. The README states the core scanner uses the official cel-go v0.32.0 runtime, and that CEL evaluated 154 candidates without proposing a suppression or falling back in five shadow runs. The reason it only shadows is stated plainly: on the locked source-disjoint split the scanner scores TP=65, FP=42, TN=503, FN=774, which is 60.75% precision, 7.75% recall, 13.74% F1 and 7.71% false-positive rate. The README says this does not pass the promotion gate, so every bundled CEL rule remains in shadow. Publishing that number is unusual and it is the single most useful thing on the page. An optional Meta-analyzer correlates and prioritizes findings and can filter them, though the README notes paired accuracy validation is still pending.
Installing cisco-ai-skill-scanner and running a first scan
The package is published on PyPI as cisco-ai-skill-scanner and requires CPython 3.11 through 3.14. The repository is built with hatchling and hatch-vcs, and its own CI pins reproducibility through uv.lock with `uv sync --frozen` plus hash-pinned requirements published per release, according to the dependency policy note in pyproject.toml. For a first look, a plain pip install is enough:
pip install cisco-ai-skill-scannerThe README links a Quick Start guide that promises a five-minute setup, and the repository ships runnable examples under examples/, including basic_scan.py, api_usage.py and batch_scanning.py. Those files are the most reliable place to copy a working invocation from, because the README itself does not print a full command line. The scanner targets standard Codex and Cursor skill layouts; for non-standard formats such as Claude Code `.claude/commands/*.md` files and flat markdown skill repositories, the README says to pass `--lenient`.
LLM analysis is opt-in and configured through environment variables. The .env.example file shows the shared prefix, with Anthropic, OpenAI, Google AI Studio, Azure OpenAI, AWS Bedrock and Google Vertex AI routes all documented in comments:
# SKILL_SCANNER_LLM_API_KEY=your_api_key
# SKILL_SCANNER_LLM_MODEL=claude-3-5-sonnet-20241022
# SKILL_SCANNER_LLM_MAX_TOKENS=16384
# SKILL_SCANNER_LLM_REASONING_EFFORT=lowNote the comment in that file: direct Google GenAI SDK requests reject the reasoning-effort control, while LiteLLM-backed Gemini requests support it. If you skip these variables entirely, the deterministic engines still run; the LLM judge and the Meta-analyzer are what you lose.
CI integration and the exit-code contract
The scanner is built to fail builds. It emits SARIF for GitHub Code Scanning, ships a reusable GitHub Actions workflow documented in docs/github-actions.md, and uses exit codes so a pipeline can stop on findings. There is also a pre-commit hook through the standard pre-commit framework, with .pre-commit-config.yaml and .pre-commit-hooks.yaml present at the repository root.
That combination is the strongest argument for the tool. A skill repository that runs a scanner on every commit and uploads SARIF gets a review surface where findings appear next to the diff. The catch is threshold tuning. Because false positives and false negatives both occur by the project's own admission, an exit-code gate set too aggressively will block legitimate skills, and one set too loosely will pass everything. The README points to docs/user-guide/custom-policy-configuration.md for presets and tuning, and to a policy quick reference for the individual knobs. Treat the default policy as a starting point to measure against your own corpus, not as a calibrated gate.
The evaluation numbers are the honest part, and also the warning
Two benchmark results are reported and they disagree sharply. On the core plus CEL development benchmark, over 5,256 malicious and 1,338 benign MaliciousSkillBench packages, the scanner raised F1 from 32.92% to 47.73% and recall from 19.88% to 31.43% against origin/main, while cutting the benign false-positive rate from 3.59% to 1.05%, at 99.16% precision. Those are development-set numbers.
On the locked source-disjoint split, the same scanner records 7.75% recall. That is the number to plan around, because it is the one the project itself treats as disqualifying for promotion. It means roughly nine out of ten malicious samples in that split were not flagged. A scanner with that recall profile is a tripwire, not a filter. It will catch some known patterns and it will miss most of what a determined author writes. The README also reports supplemental checks: identical CEL-OFF and CEL-SHADOW findings across 111 official Codex, Claude Code and Cursor skills, 0 of 339 actionable matches on the NotInject hard-negative set, 7 of 200 on HarmfulSkillBench and 76 of 263 on OpenSkillRisk. The last two are described as positive-only recall diagnostics that cannot measure precision or false-positive rate, which is the correct caveat to attach to them.
Where Skill Scanner is the wrong tool
Three cases stand out. First, if you need assurance rather than signal. The README is explicit that the scanner does not certify security and that a no-findings result does not guarantee a skill is benign. Using it to sign off a production agent deployment would misread what it produces.
Second, if your skills are not in a supported format. The scanner targets OpenAI Codex Skills and Cursor Agent Skills following the Agent Skills specification. Claude Code `.claude/commands/*.md` and flat markdown repositories need `--lenient`, and the README presents that as a concession rather than a supported path. A team whose skills live in a homegrown format should expect to write custom rules first, using the rule authoring guide for signature, YARA and Python rules.
Third, if you cannot tolerate false positives in a blocking position. At a 7.71% false-positive rate on the source-disjoint split, a gate that blocks merges will produce friction. The Meta-analyzer can filter and prioritize findings, but the README states its paired accuracy validation remains pending, so it is not yet a calibrated filter. Run it in report-only mode until you have measured it on your own skills.
Alternatives and how the approach differs
The obvious comparison is a general-purpose dependency scanner such as Snyk. The difference is the artifact. Snyk and similar tools reason about packages, versions and known vulnerability identifiers; they match your lockfile against advisory data. An Agent Skill has no versioned dependency graph to match against, and its risk lives in prose instructions and in code that may never be imported by your application. Skill Scanner instead reads the skill's own text and code, using YARA-X signatures, dataflow analysis and an optional LLM judge. Neither approach substitutes for the other, and a repository containing both application code and skills needs both kinds of tooling.
A second comparison is a plain static analyzer for Python, such as a linter with security rules. Those catch dangerous calls in code that runs. They have nothing to say about a markdown file instructing an agent to exfiltrate a file, which is the case Skill Scanner was built for. The reverse also holds: Skill Scanner's code-level analysis is narrower than a mature SAST tool's, so it is not a replacement for one.
Maintenance, licensing and upgrade cost
The last push to the default branch was on 2026-09-05, and releases 2.1.0 and 2.0.14 both landed that same day, with 2.0.13 on 2026-08-03. The repository is not archived. The project classifies itself as Development Status 4 - Beta in pyproject.toml, which is consistent with a rule set and a decision layer still under evaluation.
The licence metadata is worth reading carefully. The README badge and pyproject.toml both state Apache-2.0, and the classifiers list the OSI-approved Apache Software License, but the repository's detected licence is NOASSERTION. That usually means the LICENSE file does not match a standard template exactly. If you need the licence terms to be unambiguous for redistribution, read the LICENSE file itself rather than the badge.
Upgrade cost comes from the rule and policy surface. The CEL rules are currently in shadow and the README says this release does not promote any suppression, so a future promotion would change findings without changing your configuration. The scan policy, presets and rule packs are all configurable, which means a version bump can shift results. Pin the version in CI and re-run your own corpus before accepting a new one.
Editorial conclusion
Adopt Skill Scanner if you ship or consume Codex and Cursor skills and want signature, dataflow and optional LLM checks wired into CI through SARIF output and exit codes. Do not adopt it as a certification gate: the README states that no findings does not mean no risk, and the locked source-disjoint split reports 7.75% recall with a 7.71% false-positive rate, so bundled CEL rules stay in shadow. Before you rely on it, run the scanner against your own skill set, read the detection evaluation and rollout document for methodology, and check the scan policy configuration to see which presets and thresholds you are actually enabling.
Frequently asked questions
What is a skill check in the context of Skill Scanner?
It is a scan of an AI Agent Skill package for known risk patterns. Skill Scanner reads a skill's instructions and code using YAML and YARA-X signatures, AST and dataflow analysis, an optional LLM judge and a bounded CEL decision layer, and reports findings.
What is the purpose of scanning with Skill Scanner?
The README describes it as detecting prompt injection, data exfiltration and malicious code patterns in Agent Skills. It is a detection aid, not a certification: the README states that a scan returning no findings does not guarantee a skill is free of all threats.
How do I install cisco-ai-skill-scanner?
It is published on PyPI as cisco-ai-skill-scanner and requires CPython 3.11 through 3.14, so a pip install of that package name is the documented route. The repository also ships runnable examples under examples/ and a Quick Start guide linked from the README.
Does Skill Scanner work on Claude Code skills?
The README says the scanner supports OpenAI Codex Skills and Cursor Agent Skills formats following the Agent Skills specification, and that non-standard formats such as Claude Code `.claude/commands/*.md` and flat markdown skill repositories are scanned with `--lenient`.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/cisco-ai-defense-skill-scanner)