# AgentSeal: a local security scanner for AI agent skills and MCP configs

> AgentSeal bundles four commands that check agent skill files, MCP server configs and live tool descriptions for poisoning and exfiltration paths. Here is what each one actually does, how to install it, and where it stops being useful.

**getagentseal/agentseal** — Security toolkit for AI agents. Scan your machine for dangerous skills and MCP configs, monitor for supply chain attacks, test prompt injection resistance, and audit live MCP servers for tool poisoning.

- Repository: https://github.com/getagentseal/agentseal
- Website: https://agentseal.org
- Stars: 377 · Forks: 69
- Language: Python
- License: NOASSERTION
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/getagentseal-agentseal

## The problem AgentSeal targets: instructions your agent reads but you never see

An MCP server hands an agent access to local files, databases, APIs and credentials. The agent also reads each tool's description text, and that text is not shown to the user in most clients. The README states the risk plainly: tool descriptions can contain hidden instructions that the agent follows but the user never sees. The same applies to skill files, which are instruction documents loaded into a coding agent's context.

The repository also points at a second, quieter problem: an npm install or pip install can silently modify your agent configs. Nothing in the install output tells you a config file changed. AgentSeal is aimed at engineers who already run several agents on one machine, which is why the README lists 28 supported agents including Claude Code, Cursor, Windsurf, VS Code, Gemini CLI, Codex CLI, Cline, Aider and Zed. If you run one agent and never install third-party MCP servers, the surface this covers is small.

## How guard works: six detection stages and a baseline diff

guard is the command that needs no API key and makes no network calls, according to the README. It walks the agent configuration paths on your machine and runs a six-stage pipeline on each file it finds.

Stage one is pattern signatures for known malicious shapes: credential access, exfiltration URLs, shell commands. Stage two deobfuscates, decoding Unicode tags, Base64, BiDi overrides, zero-width characters and TR39 confusables, which is where a payload hidden as invisible characters gets turned into something the later stages can read. Stage three is semantic analysis using embedding similarity with MiniLM-L6-v2, intended to catch rephrased attacks that slip past literal patterns. Stage four is baseline tracking: SHA-256 hashes of configs are compared against your last scan, which the README calls rug-pull detection, meaning a server that behaved on day one and changed later. Stage five enriches findings with trust scores from the MCP Security Registry, described as covering 6,600+ servers. Stage six applies your own YAML rules.

The design choice worth noting is that stages one through three are local heuristics. Semantic similarity against a MiniLM embedding will surface paraphrases a regex misses, but it also produces similarity judgements rather than proof, and the README does not document a false-positive rate for that stage. Treat a semantic hit as a reason to open the file, not as a verdict.

## Installing AgentSeal and running a first guard scan

AgentSeal ships on PyPI and npm. The README gives both install paths and then a single command that needs no key.

```bash
pip install agentseal    # or: npm install agentseal
agentseal guard          # scan your machine - no API key needed
```

The guard run prints findings for dangerous skill files, poisoned MCP server configs and data exfiltration paths across the agents it detects. Because this first run establishes the hash baseline, it is the one scan whose output you should actually read rather than skim.

Once the baseline exists, you can generate a project policy file and change the output format for CI use:

```bash
agentseal guard init             # generate .agentseal.yaml project policy
agentseal guard --output sarif   # SARIF for GitHub Security tab
agentseal guard --output json    # machine-readable output
agentseal guard --no-diff        # skip baseline delta section
agentseal guard test             # validate your custom rules
```

The .agentseal.yaml file is where org-specific rules live, and guard test validates them before you depend on them. If you want continuous watching rather than a one-shot scan, shield is a separate install with extra dependencies:

```bash
pip install agentseal[shield]   # includes watchdog + desktop notification deps
agentseal shield
```

Shield monitors the same paths guard scans, raises desktop notifications and quarantines files with detected payloads. The README does not document an undo or restore command for quarantined files, so check where your files go before running it on a working machine.

## Testing a prompt with scan, and the canary trick behind the trust score

scan is the part that needs a model. It runs a system prompt against 225 adversarial attack probes, broken down in the README as 82 extraction techniques, 143 injection techniques and 8 adaptive mutation transforms, and returns a trust score from 0 to 100.

The detection mechanism is deterministic by construction. Injection probes embed a unique canary string, with the README's example being SEAL_A1B2C3D4_CONFIRMED. If the canary shows up in the response, the probe leaked. Extraction probes use n-gram matching against the ground truth prompt. There is no LLM judge, so the same input produces the same result every time. That is a real advantage over rubric-based prompt testing, which drifts between runs and makes score comparisons across commits unreliable.

The trade-off is coverage. A canary only fires when the model reproduces that exact token, so a model that leaks the substance of a prompt without emitting the canary is scored as holding. The score bands are 85 to 100 Excellent, 70 to 84 High, 50 to 69 Medium, 30 to 49 Low, and 0 to 29 Critical. You can point it at a cloud model, a local Ollama model, or any HTTP endpoint:

```bash
agentseal scan --prompt "You are a helpful assistant..." --model gpt-4o
agentseal scan --prompt "You are a helpful assistant..." --model ollama/llama3.1:8b
agentseal scan --url http://localhost:8080/chat
agentseal scan --file ./prompt.txt --model gpt-4o --min-score 75
```

The --min-score flag is the CI hook: exit code 1 when the trust score falls below the threshold. For a system prompt that ships with a product, that is a more useful gate than a manual review, provided you accept that the gate measures canary leakage and n-gram overlap rather than whether the prompt is safe in a broader sense.

## Auditing a live MCP server with scan-mcp

scan-mcp is the command that inspects a server rather than your disk. It connects over stdio or SSE, enumerates every tool the server exposes, and runs each tool description through pattern matching, deobfuscation, semantic similarity and optional LLM classification, then reports a trust score per server.

The README gives both connection forms:

```bash
agentseal scan-mcp --server npx @modelcontextprotocol/server-filesystem /tmp
agentseal scan-mcp --sse http://localhost:3001/sse
```

The stdio form launches the server command you pass, so you are executing the same npx invocation your agent would. That is deliberate, since it audits the server as configured, but it also means scan-mcp is not a sandbox: anything the server does on startup happens on your machine. Run it against servers you already intend to use, not as a way to safely preview an unknown package. The SSE form only needs a URL, which is the lower-risk option when a server is already running in a container or on a remote host.

## Where AgentSeal is the wrong tool

Guard reads agent configuration paths. It is not a general-purpose scanner for your application code, your dependencies or your container images, and the README does not present it as one. If a supply chain question is about a Python package you import rather than an agent config that package modifies, guard will not answer it.

The scan command has a narrower ceiling than the probe count suggests. Its verdicts come from canary strings and n-gram matching, so it measures whether a specific prompt leaked under specific probes. It says nothing about tool-call abuse, indirect injection through retrieved documents, or multi-turn attacks, none of which the README lists among the 225 probes. A high trust score means the prompt resisted this probe set, not that the agent is safe.

There is also a coverage gap on the guard side. The supported agent list is long but finite, and the pipeline only sees files inside the paths it knows about. An agent or MCP client not on that list can hold a poisoned config that guard never opens. The README does not document a way to add arbitrary scan paths, so the custom YAML rules in .agentseal.yaml extend detection rather than reach.

## How AgentSeal differs from generic static analysis and prompt-testing suites

The nearest comparison is a general SAST tool such as Semgrep. Semgrep parses source files and matches rules against code structure across a repository you point it at. AgentSeal does not parse your code; it reads the instruction and configuration files that agents consume, and its stages are tuned for that content: Unicode tag decoding, BiDi override detection and TR39 confusables are about hiding text in a description, not about finding a tainted variable. If your concern is a vulnerable function in your own service, Semgrep is the right shape of tool. If your concern is a tool description that tells the agent to read ~/.ssh and post it somewhere, Semgrep has no reason to look there.

The other comparison is prompt-testing frameworks that use a model as the judge. Those can evaluate open-ended behaviour and multi-turn flows, which AgentSeal's canary approach cannot. They also produce scores that move between runs for reasons unrelated to the prompt. AgentSeal's determinism is the point of the design, and it is the reason the trust score works as a CI gate. The cost is that the score is a lower bound on exposure, not an estimate of it.

One more difference sits in the licence. The repository badge reads FSL-1.1-Apache-2.0 while the repository metadata reports NOASSERTION, so the two disagree. The Functional Source License is not an OSI-approved open source licence at the time of writing; it typically restricts competing use and converts to a stated open licence after a set period. Read the LICENSE file in the repository root before you build a product around the tool, and if the distinction matters to your legal team, treat the badge and the metadata as two claims that need reconciling.

## Maintenance, upgrade cost and what to check before adopting

The repository is not archived. The most recent push recorded is 2026-06-11, which is more than three months before today, and the latest release listed is v0.8.0 (Guard) from 2026-03-25, with v0.6.2 before it in March. There is no published cadence, and the gap between the last release and the last push means fixes may be landing on main without a tagged version. If you pin AgentSeal in CI, pin a released version rather than tracking main.

The upgrade cost is concentrated in two places. First, the baseline: guard compares hashes against your previous scan, so upgrading mid-stream or rescanning after a large config change resets what counts as new. Second, the probe set: scan reports a trust score against 225 probes, and if that set grows in a later release, scores computed before and after are not directly comparable. Record the AgentSeal version alongside any trust score you store.

The extra install for shield pulls in watchdog and desktop notification dependencies, so a headless CI runner is a poor place for it. Keep shield on developer machines and guard in CI. On the licence question, do not treat the FSL-1.1-Apache-2.0 badge as equivalent to Apache-2.0; check the LICENSE file and the conversion terms yourself.

## Conclusion

Adopt AgentSeal if you run multiple AI coding agents or MCP servers on a developer machine and want a local, no-API-key check of skill files and configs before trusting them. Skip it if your only need is static analysis of your own application source, since guard inspects agent configuration paths rather than your codebase. Before relying on it, run agentseal guard once and read the baseline it writes, because every later scan is a diff against that first run and a stale baseline makes the rug-pull stage less meaningful.

## FAQ

### Does AgentSeal need an API key to scan my machine?

No. The README states that agentseal guard needs no API key and makes no network calls, and that everything runs locally. Only the scan command needs a model, and it can run free against a local Ollama model such as ollama/llama3.1:8b.

### Which AI agents does AgentSeal guard scan?

The README lists Claude Code, Claude Desktop, Cursor, Windsurf, VS Code, Gemini CLI, Codex CLI, Cline, Roo Code, Kilo Code, Copilot CLI, Aider, Continue, Zed, Amp, Amazon Q, Junie, Goose, Kiro, OpenCode, OpenClaw, Crush, Qwen Code, Grok CLI, Visual Studio, Kimi CLI, Trae and MaxClaw.

### How does AgentSeal decide whether a prompt leaked?

Injection probes embed a unique canary string, and the probe counts as leaked if that canary appears in the response. Extraction probes use n-gram matching against the ground truth prompt. There is no LLM judge, so the same input produces the same result each time.

### Can AgentSeal audit an MCP server that is already running?

Yes. The scan-mcp command accepts an SSE endpoint such as http://localhost:3001/sse, or a stdio server command, and enumerates every tool before scoring each description for poisoning.

### What licence does AgentSeal use?

The repository badge reads FSL-1.1-Apache-2.0, while the repository metadata reports NOASSERTION, so the two do not agree. The LICENSE file in the repository root is the document to read.

## Sources

- [getagentseal/agentseal on GitHub](https://github.com/getagentseal/agentseal)
- [Issues](https://github.com/getagentseal/agentseal/issues)
- [Project website](https://agentseal.org)
- [README](https://github.com/getagentseal/agentseal/blob/main/README.md)
- [Releases](https://github.com/getagentseal/agentseal/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/getagentseal-agentseal
