Model or dataset
mattzcarey/shippie avatar
mattzcarey/shippie

mattzcarey/shippie: an extendable code review agent that runs on flue and pi

extendable code review and QA agent 🚢

2,503 stars243 forksTypeScriptMIT

At a glance

What is it?
Shippie wraps an agent loop around your git diff and posts review comments from a GitHub Action, a local command, or a comment trigger. It is provider-agnostic, MCP-capable, and still at 0.21.x.
Who is it for?
Adopt shippie if you already run pull-request CI and want review comments that can read the surrounding codebase rather than only the patch, and if you are comfortable pinning a 0.21.x line whose framework dependencies are still on 1.0.0-beta tags. Skip it if you need a stable, audited release cadence, if your CI cannot give a full checkout with fetch-depth: 0, or if you will not fund a provider API key.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 2 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What shippie reviews that a plain diff bot misses

Most automated review tools see a patch and nothing else. Shippie's premise is that the agent should be able to walk the repository before it comments. The README describes an agent loop that reads your diff, explores the codebase with real developer tools, and posts focused review comments, naming exposed secrets, slow or inefficient code, and potential bugs or unhandled edge cases as the categories it targets. That ordering matters: the diff is the entry point, not the whole input.

The intended user is a team already running pull-request CI on GitHub or GitLab. Shippie is not a linter and does not try to be one. It is also not a general-purpose coding agent you chat with; the README frames it as a prebuilt review workflow rather than a bespoke CLI. If you want an agent you drive by hand, this is the wrong shape. If you want a review step that appears on every pull request without anyone invoking it, that is the shape it has.

The flue and pi agent loop, and why the framework choice is the design

The mechanism is stated plainly in the ethos section: the agent loop runs on flue plus pi, and shippie functions as a human code reviewer using flue's built-in tools instead of a hand-rolled tool registry. That is the architectural decision worth understanding. Rather than defining its own set of tools for reading files or searching the tree, shippie inherits them from the framework, and the review workflow is expressed as a flue workflow. The package scripts confirm this: review and qa are both flue run invocations, and the build step produces a Node server at dist/server.mjs.

Two consequences follow. First, shippie's capability ceiling moves with flue, which is why the dependency list carries @flue/runtime and @flue/github at 1.0.0-beta tags. Second, the provider layer is separate from the loop. The .env.example lists Anthropic, OpenAI, OpenRouter, and Cloudflare Workers AI credentials, with SHIPPIE_MODEL defaulting to anthropic/claude-sonnet-4-6, and CLOUDFLARE_API_KEY accepting either a Cloudflare API token or a wrangler login OAuth token with AI access. Swapping providers is an environment change, not a code change.

The MCP angle is the other half. Shippie can act as a Model Context Protocol client to reach external tools such as browser automation, observability, and documentation, configured through the MCP_SERVERS action input. That is where a review agent stops being a diff reader: it can query a running system or a docs endpoint mid-review. The README documents the client side; it does not claim shippie exposes an MCP server of its own.

Installing shippie as a GitHub Action or running it locally

The README gives two entry points. The fastest is the scaffolder, which writes the workflow file for you; you then add a provider API key as a repository secret. The action needs a full checkout, so fetch-depth: 0 is not optional, and it needs pull-requests: write to post comments.

bash
npx shippie init

If you would rather write the workflow by hand, the README publishes it in full. This is the shape to copy:

yaml
name: Shippie

on:
  pull_request:

permissions:
  pull-requests: write
  contents: read

jobs:
  review:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
        with:
          fetch-depth: 0
      - uses: mattzcarey/shippie@v0
        env:
          GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
          ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}

For a first run without touching CI, use local mode. The README states that it reviews your staged changes via git diff --cached and writes results to .shippie/review/local_*.md. Stage something first, then run the command:

bash
npx shippie review

What you should see is a markdown file under .shippie/review/ rather than comments on a pull request. That is the cheapest way to judge whether the agent's comments are useful on your codebase before you give it write access to pull requests. If you are building from source instead, the README targets Node >= 22.19 with npm, and the sequence is npm install, copy .env.example to .env, set the provider key matching your model, then npm run review. There is also a Dockerfile for the QA workflow that installs system Chromium and ffmpeg and sets CHROME_BIN=/usr/bin/chromium, with the comment that the same image runs locally and in CI.

Where shippie gets in the way

The version number is the first honest signal. The latest release in the repository is v0.21.2, and the core runtime dependencies are pinned to 1.0.0-beta versions of the flue packages. Pre-1.0 software with beta framework dependencies is not a stability claim, it is the opposite. Expect the workflow shape and the action inputs to move.

The second constraint is environmental. A review agent that explores the codebase needs the codebase, so shallow checkouts break the premise. The README is explicit that the action requires fetch-depth: 0. Teams that clone shallow for speed, or that run review on a merge queue with a synthetic ref, will need to change that before shippie is useful.

Third, cost and telemetry. Every review is model calls against your provider key, and the README's .env.example shows SHIPPIE_TELEMETRY defaulting to true, with the comment that setting it to false opts out of anonymous usage telemetry. That is a default you should decide about deliberately rather than inherit. There is also retry tuning exposed through PI_AI_RETRY_ATTEMPTS, PI_AI_RETRY_BASE_MS, and PI_AI_DISABLE_RETRY, with the example describing exponential backoff delays of roughly 500, 1500, 4500, and 13500 milliseconds. If your provider bills retries, that is real spend.

Finally, scope. Shippie reviews and posts comments. It does not block merges by itself, and the README does not document a pass/fail gate you can require in branch protection. If your policy is that a review agent must be able to fail a check, that has to be built around it.

Shippie against a rule-based linter plus a hosted review bot

The obvious alternative is the combination most teams already have: a linter and static analysis in CI, plus a hosted review bot that comments on pull requests. The difference in approach is what the comment is derived from. Linters evaluate syntax and known patterns against rules someone wrote down. Hosted review bots typically send the diff, and sometimes a slice of surrounding files, to a model with a fixed prompt.

Shippie's bet is that an agent with tools produces better comments than a model with a bigger prompt, because the agent can go look. It can read the file the changed function calls into, or query an observability tool through MCP, before deciding whether the change is a bug. That is a genuine difference in kind, not a tuning difference.

The trade-off is predictability. A linter's output is deterministic and its cost is CPU. Shippie's output depends on the model, the thinking level, and what the agent happened to explore, and its cost is tokens. The README exposes SHIPPIE_THINKING_LEVEL with a default of medium, which is an admission that the depth of exploration is a dial you will end up tuning. If your team's complaint about AI review is inconsistent output, an agent loop does not obviously fix that; it may make it more variable. If your complaint is that the bot only ever comments on the patch in front of it, this is aimed squarely at you.

Maintenance, licence, and what an upgrade actually costs

The repository is not archived, and the last push was on 2026-09-10, so work is landing. The release history shows v0.21.0, v0.21.1, and v0.21.2 clustered on 2026-06-22 and 2026-06-23, and the project uses changesets for versioning, with the release script running a build and then changeset publish. In practice that means upgrade notes arrive as changeset entries and a CHANGELOG.md rather than as a maintained migration guide; the README points to docs/setup.md for setup but does not document a rollback path for a bad upgrade.

Because the action is referenced as mattzcarey/shippie@v0, you are tracking a moving major tag unless you pin a specific version. That is a choice worth making explicitly: the tag gets you fixes, a pinned SHA gets you reproducibility. The same applies to the npm package, where npx shippie review resolves to whatever the latest published 0.21.x is.

Licensing is straightforward: the package.json declares MIT, and the repository carries a LICENSE file. MIT permits commercial use and modification, and it comes with no warranty. That is a statement about the licence text, not advice about your situation; if your organisation has rules about model-generated content in review comments or about telemetry defaults, those are separate questions the licence does not answer.

Editorial conclusion

Adopt shippie if you already run pull-request CI and want review comments that can read the surrounding codebase rather than only the patch, and if you are comfortable pinning a 0.21.x line whose framework dependencies are still on 1.0.0-beta tags. Skip it if you need a stable, audited release cadence, if your CI cannot give a full checkout with fetch-depth: 0, or if you will not fund a provider API key. Before rolling it out, run npx shippie review against your staged changes and read .shippie/review/local_*.md to see what the agent flags on your own code, then check docs/action-options.md for the MODEL and IGNORE inputs you will need to keep noise down.

Frequently asked questions

What is shippie?

Shippie is an extendable code review agent that runs an agent loop over your diff, explores the codebase with developer tools, and posts focused review comments. The README lists exposed secrets, slow or inefficient code, and potential bugs or unhandled edge cases among the issues it targets.

How do I install shippie in a GitHub Actions workflow?

Run npx shippie init to scaffold the workflow, then add your provider API key as a repository secret. The action requires a full checkout with fetch-depth: 0 and pull-requests: write permissions.

Can I run shippie locally without a server?

Yes. The README states that local mode reviews your staged changes via git diff --cached and writes results to .shippie/review/local_*.md, with no server involved.

Which model providers does shippie support?

The .env.example lists Anthropic, OpenAI, OpenRouter, and Cloudflare Workers AI credentials, with SHIPPIE_MODEL defaulting to anthropic/claude-sonnet-4-6. Cloudflare Workers AI also accepts a wrangler login OAuth token with AI access.

Does shippie send telemetry by default?

The .env.example shows SHIPPIE_TELEMETRY defaulting to true, with the comment that setting it to false opts out of anonymous usage telemetry.

Official sources

  1. License: MIT
  2. mattzcarey/shippie on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/mattzcarey-shippie.svg)](https://hysenlabs.com/projects/mattzcarey-shippie)