# Ship Safe ranks its own passes so the cheap heuristic can never overrule the data-flow trace

> A local scanner for AI-written software with two layers: a deterministic engine that finds issues across application code, agents, MCP configs, dependencies, pipelines and secrets, and an investigation layer that traces each finding and reports four verdicts with the lines it read. No key is needed for the core, three named features are not local, and the project ships benchmarks for its own false positives.

**asamassekou10/ship-safe** — The independent security agent for AI-written software. Finds issues, investigates whether they are real, and shows you the evidence. Deterministic core, no API key needed, JSON and SARIF output.

- Repository: https://github.com/asamassekou10/ship-safe
- Website: https://shipsafe.sh
- Stars: 848 · Forks: 113
- Language: JavaScript
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/asamassekou10-ship-safe

## Two layers, and the cheaper pass is never allowed to win

The tool runs in two layers and the second one is the interesting half.

The first is a deterministic engine that finds candidate issues across application code, AI agents, MCP configuration, prompts, dependencies, pipelines, secrets and cloud-adjacent settings. That layer is described as fast, repeatable and benchmarked, and it is the sensor.

The second decides what a finding is worth. It traces a value to the sink it reaches, looks for controls elsewhere in the project that a rule assumed were missing, joins attack chains spread across configuration no single file contains, and will, if asked, probe a leaked key against the provider that issued it. Every conclusion carries both the pass that reached it and the lines that were read.

The ranking rule is stated explicitly, and it is the design's core: a traced data path outranks a model's reading of the same file, and a probe that actually authenticated outranks both. Separate passes are ranked so that a cheaper one can never overturn a more expensive one. A heuristic match is a candidate; it becomes a conclusion only when something stronger agrees.

Output resolves to four verdicts, confirmed, likely, unresolved and refuted, each citing the lines it was concluded from. The stated benefit is that you can disagree with a step rather than with a severity label. The shape of a real finding, from a documented run against a deliberately vulnerable application, looks like this:

```
CONFIRMED - traced end to end (10)

    NoSQL Injection via $where [high]
    app/data/allocations-dao.js:78  NOSQL_INJECTION_WHERE
    why: threshold is assigned from the HTTP request and reaches the sink without validation on that path.
    decided by: dataflow
      1. value reaches NOSQL_INJECTION_WHERE here  app/data/allocations-dao.js:78
      2. getByUserIdAndThreshold is called here with threshold  app/routes/allocations.js:23
      3. threshold is assigned here  app/routes/allocations.js:20
    fix: Replace $where with standard MongoDB operators ($eq, $gt, $regex, etc.)
```

Three numbered steps and one sink, with the path crossing three directories between assignment and use.

## No key for scanning, and three named features that are not local

The headline is a single command with no signup and no API key required for scanning, and core checks are said to work offline.

There is also a flag that guarantees a fully local scan, which is the clearest signal that the default is not fully local. Three features are named as sending bounded context to whichever provider you have configured: provider-backed classification, deep analysis, and a red-team mode. Before sending, the tool attempts to mask credentials, and the wording is best-effort. The exact boundaries and context limits are documented on a separate page rather than in the main document.

So the claim decomposes into two parts. The deterministic engine and the data-flow investigation need nothing from you. The parts that ask a model to read or classify something go to a provider you nominate. For a tool whose entire pitch is reviewing code you wrote with an AI assistant, that distinction matters twice over: it is where your source goes, and it is where your provider's data retention terms apply.

The audit command widens the surface in a different way. It is described as a full audit covering secrets, thirty agent configurations, dependencies and a remediation plan, which is a much broader scope than the trace-oriented investigate command and a much larger run.

Investigate itself has three flags: the default, a flag that also details the unresolved and refuted findings rather than summarising them, and a verification flag that probes leaked keys against their providers.

## The CI baseline stores hashed finding identities and no raw secrets

The pull request workflow is the part of this design that most obviously came from having been used on a real repository.

The problem it solves: a scanner run on a pull request will report every pre-existing problem in the repository, so a team either blocks on debt they did not create or turns the check off. The solution is a two-scan comparison, a baseline on a trusted revision and the head scan compared against it:

```bash
# On the trusted base revision
npx ship-safe ci . --fail-on none --no-deps \
  --write-baseline-report /tmp/ship-safe-base.json

# On the pull request head
npx ship-safe ci . --base-report /tmp/ship-safe-base.json --fail-on high
```

The baseline is deliberately not a findings dump. It holds hashed finding identities, relative paths and rule metadata, and the document states in as many words that it does not store raw matched secrets. That matters because a baseline file written to a temporary path in a pipeline is exactly the kind of artefact that gets archived, cached or uploaded by something else, and one containing live credentials would be a new leak created by the tool meant to find them.

Classification against a baseline has four outcomes: introduced, resolved, unchanged, or uncertain. The uncertain bucket is the honest one, produced by ambiguous matches, and those are shown without blocking the pull request. The two scans use different failure thresholds, none on the base and high on the head, so the base scan records rather than judges.

Findings can also be uploaded as SARIF into GitHub code scanning, which is what puts the results next to the diff rather than in a separate dashboard.

## There is a command for the folder you are about to open with an agent

The most distinctive command is the one that examines a folder you have not trusted yet, and the framing in the document is a warning: before you open an unfamiliar directory with an agent, find out what runs on open. Pointed at a downloads folder, it asks what that directory would cause to execute. It takes a JSON output flag for the cases where you want to feed the answer into something else.

The related command is `capabilities`, and it exists to answer a question a coding agent cannot answer about itself. An agent reviewing your repository cannot read your MCP server configuration, cannot enumerate the permissions it was launched with, and is itself the actor whose reach is in question. This command reads all of that from outside and reports combinations that are unremarkable on their own and dangerous together. A documented critical finding reads like this, and the pattern is worth more than the specific rule:

```
  CRITICAL  Repository-controlled instructions reach an unattended write capability
    1. CLAUDE.md is read as instructions and can be changed by anyone who lands a commit
       CLAUDE.md:1
    2. Claude Code runs without per-action approval
       .claude/settings.json:2
```

Neither fact is alarming alone. A project instruction file is normal, and running without per-action approval is a common setting. Together they mean anyone who can land a commit can change the instructions given to an agent that can write without asking.

This is the same argument the project makes about reviewing your own code, extended from code to configuration. A separate reviewer with a separate method is the whole premise, and this command is where it stops being a slogan.

## Fixing asks before writing and keeps the change reversible

The remediation path is described as a five-step process, and the fourth step is the one that distinguishes it from a tool that edits your files.

Scanning is local and uses targeted checks, skipping the ones that do not apply. Each finding is then investigated in separate passes under the ranking rule already described. The evidence is shown with the lines it came from. And then, to fix: the agent proposes a plan and a diff, asks before writing anything, verifies the result, and keeps the changes reversible.

Three properties in one sentence. It does not start editing on its own initiative. It checks its own work before declaring the problem handled, which is the same standard it applies to the code it scanned. And it leaves a way back.

The scope of what it looks at is broad, and the table divides it into six areas. Artificial intelligence and language model security covers prompt injection, agent hijacking, excessive agency, memory poisoning, retrieval poisoning and unsafe tool calls. Agent and MCP configuration covers over-broad tool permissions, poisoned registries, untrusted transports and dangerous allowlists. Application security is the conventional set, including injection in both its SQL and NoSQL forms, cross-site scripting, server-side request forgery, authentication bypass and path traversal. Secrets and compliance covers keys, tokens, credentials, personal data and secrets in git history. Supply chain covers typosquatting, dependency confusion, risky install scripts and unpinned AI actions. And pipelines cover poisoning, unpinned GitHub Actions, secret logging and unsafe triggers.

That last area is the one most tools skip, and for AI-written software it is not decorative, since an unpinned action referenced by a workflow is remote code execution with a version range instead of a version.

## The benchmarks include a suite for its own false positives

The repository root carries seven benchmark entry points, and three of them are about judging the tool rather than finding bugs.

There is a corpus suite, a verdicts suite, an applications suite, and a suite that runs release evidence checks. Two of them have the shape you want to see in a scanning project: a false-positive suite, and a verdicts suite with a write mode. The write modes matter because they mean expected output can be regenerated from a run rather than only checked against a frozen expectation, which is what lets a rule change be reviewed as a diff of expected verdicts.

There is also a script that checks the roadmap. The package scripts include a roadmap check that runs a script against the roadmap file, so the state of the project's own promises is verified somewhere in the same pipeline that runs the tests.

Tests use the test runner built into the runtime, invoked over the test files in one directory, with no external framework in the dependency list. Linting is scoped to the command line directory, and the package entry points are a module file for library use and a binary script under the same directory.

Version history shows a deliberate major bump: 9.9.0, then 10.0.0 on 2026-09-05, then 10.1.0 on 2026-09-21 with a release note about attestation precision. The manifest version matches the newest tag, and the default branch was pushed the same day as that tag.

## It ships as an agent plugin, a GitHub Action, and a printed approval command

The distribution surface is broader than a command line tool, and each piece targets a different way code gets written.

At the root there is a plugin directory aimed at a specific coding agent, a skills directory, a directory of text snippets, and a JSON descriptor for yet another agent runtime. There is a workflow definition file, which makes the scanner usable as a GitHub Action rather than only as something a developer types. A directory named for virtual machines sits beside a release assets directory.

There is also an ignore file, which is the right sign: a scanner with no way to exclude a generated directory trains people to run it with reporting switched off instead.

One script in the manifest deserves a second look. It queries the forge's own API for workflow runs in a state that requires manual approval, and prints a formatted line per run with the timestamp, the workflow name, the branch, the run identifier, and the exact command needed to approve that run. It does not approve anything itself. It builds the command and leaves it on the screen.

That is the same posture as the fix path and the risk classification: the tool does as much as it can safely do on its own, hands the irreversible step back as a copyable command, and keeps a record. Whether that is the right balance for an automated review tool is a question each team has to answer, but it is at least a deliberate position rather than an accident.

## Conclusion

Ship Safe is worth taking seriously on one narrow point: it refuses to let a model's reading of a file outrank a traced data path through it, and it shows you the numbered trace so you can argue with a step instead of a severity. The pull request baseline is the other feature that will change how a team uses it, because hashing finding identities and excluding raw secrets from the artifact is what makes it safe to write to a temporary path in a pipeline. Before adopting it, understand three limits. Three named features, namely provider-backed classification, deep analysis and a red-team mode, send bounded context to a provider you choose, with credential masking described as best-effort, so the offline guarantee applies to the core and not the product. The trust and capabilities commands are the ones that answer a question no other tool in this category asks, and they are the reason to try it. And the benchmark directory includes a suite for its own false positives, which is the number to look for rather than the detection count.

## FAQ

### What does Ship Safe do and does it need an API key?

It is a local security scanner for AI-written software with two layers: a deterministic engine that finds candidate issues across application code, agents, MCP configuration, dependencies, pipelines and secrets, and an investigation layer that traces each finding and reports confirmed, likely, unresolved or refuted with the lines it read. No API key is required for scanning and core checks work offline, but provider-backed classification, deep analysis and a red-team mode send bounded context to a provider you configure.

### How does Ship Safe decide whether a finding is real?

Investigations run as separate passes ranked by cost, so a cheaper one never overturns a more expensive one. A traced data path outranks a model's reading of the same file, and a probe that actually authenticated against a provider outranks both. Findings resolve to confirmed, likely, unresolved or refuted, each citing the lines the conclusion came from, and each carries the name of the pass that reached it.

### Does Ship Safe block a pull request on pre-existing problems?

No. The CI mode writes a baseline report on a trusted base revision and compares the pull request head scan against it, classifying findings as introduced, resolved, unchanged or uncertain. The base report holds hashed finding identities, relative paths and rule metadata, and does not store raw matched secrets. Ambiguous matches are shown but do not block, and results can be uploaded as SARIF into GitHub code scanning.

### What does the Ship Safe trust command do?

It is meant for a folder you have not opened yet. Pointed at a directory such as one in your downloads, it reports what would run if you opened it with an agent. A JSON output flag is available. The related capabilities command reads your MCP server configuration and launch permissions from outside the agent, and reports combinations that are unremarkable individually and dangerous together, such as repository-controlled instructions reaching a write capability that runs without per-action approval.

### Will Ship Safe edit my files without asking?

The described fix path has the agent propose a plan and a diff, ask before writing anything, verify the result, and keep the changes reversible. Scanning and investigation run locally and do not modify the repository. Repository-level exclusions are handled through an ignore file at the root of the project.

## Sources

- [asamassekou10/ship-safe on GitHub](https://github.com/asamassekou10/ship-safe)
- [License: MIT](https://github.com/asamassekou10/ship-safe/blob/main/LICENSE)
- [Project website](https://shipsafe.sh)
- [README](https://github.com/asamassekou10/ship-safe/blob/main/README.md)
- [Releases](https://github.com/asamassekou10/ship-safe/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/asamassekou10-ship-safe
