Model or dataset
taovc/pr-cockpit avatar
taovc/pr-cockpit

PR Cockpit ships one rolling nightly and nothing is code signed

Local AI PR workbench — Claude/Codex review in isolated git worktrees, human-gated GitHub comments

494 stars2 forksTypeScriptMIT

At a glance

What is it?
A local review workbench that pulls a repository's pull request queue into a desktop app, has Claude or Codex review each one in a read only worktree, and refuses to post anything until a human ticks a box. The distribution is a rolling pre-release with no versioned tags, the packages are unsigned on all three platforms, and the agent SDK is pinned to one exact version while everything around it floats.
Who is it for?
PR Cockpit is a careful design for the specific failure everyone has with agent assisted review, which is publishing a confident wrong comment under your own name. Its read only review agents, its per finding gate and its transactional publishing are the parts worth copying even if you use something else.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 11 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The release list holds one nightly from July and nothing else

Distribution is a rolling pre-release. The readme says builds are published as a nightly channel and rebuilt on every push to the main branch, and the download links point at the releases page rather than at a version.

The release metadata tells a different story. There is exactly one release, named for the nightly channel, dated 2026-07-01. The branch itself was last pushed on 2026-09-22, roughly eleven weeks later.

So either the pre-release is being re-published under the same name and the metadata reflects the first publication, or the nightly has not been rebuilt since July while the readme promises it is rebuilt on every push. Nothing in the visible material distinguishes those two cases, and the difference matters: a rolling channel that stopped updating months ago is a versioned release wearing a moving label.

There are no versioned releases to fall back on either. The package manifest sits at 0.1.0, is marked private, and its main entry point is the Electron main module, so the version you would pin from the manifest is not the version the download link gives you.

Three artifacts are published, one per platform: a disk image for Apple Silicon Macs, an installer for 64 bit Windows, and an AppImage for x86_64 Linux, each carrying a version in the filename. Intel Macs are explicitly not covered and are directed to building from source instead, so a large share of Mac users get a different experience from everyone else with no prebuilt artefact at all.

Nothing is signed, and each platform needs its own bypass

The readme states that none of the packages are code signed, so every platform shows a warning on first launch, and then gives one workaround per platform.

On macOS the system reports the application as damaged or from an unidentified developer, and the fix is to clear the quarantine attribute once:

bash
xattr -dr com.apple.quarantine "/Applications/PR Cockpit.app"

There is a second route for people who would rather not run that, a right click in Finder followed by Open and Open again in the dialog. Both achieve the same thing, which is telling the operating system to trust an unsigned artefact it would otherwise refuse.

On Windows the SmartScreen dialog appears with the standard protected your PC wording, and the documented path is More info and then Run anyway. On Linux there is no signing warning at all, only a permission problem, since the AppImage has to be made executable before it will start:

bash
chmod +x pr-cockpit-*-x86_64.AppImage

The three workarounds are unequal in risk. Clearing a quarantine flag is a one line command with a path in it, and it is permanent for that application. Run anyway is a click in a dialog whose whole purpose is to make you think twice. And the Linux step changes no trust relationship at all, because AppImages have never been signed by convention.

Whatever the route, the prerequisite list applies to all of them: a completed GitHub CLI authentication and a login for either Claude or Codex, before the application can do anything at all.

Review agents are read only, enforced at the tool layer

The safety model has two tiers, and the first one is unusually concrete.

Review agents are read only, and that is enforced by tool level blocking rather than by instruction. Git writes, file edits, network access and dangerous commands are all blocked at the tool boundary, and there is an operating contract and skill linting on top of that. The distinction matters: a system prompt telling a model not to write files is advice, while a tool layer that refuses the call is a boundary.

The second tier is structural. Anything that needs write access runs inside an isolated git worktree, and the operations that reach outside the machine, pushing, opening a pull request and dangerous commands, each require an explicit action in the interface or a toggle switched on deliberately.

Three consistency rules are stated alongside. Git operations against the same repository are serialised, so two tasks cannot interleave writes into one working copy. Findings are transactional, so a partially posted review is not a state the system can rest in. And a deleted task cleans up its worktree, with a restart recovering or stopping interrupted work.

The automation that would remove the human from the loop is described plainly as high risk and disabled by default. When enabled, it reuses the same review, post, fix and push endpoints as the manual path, driven by a server side poller rather than a separate implementation, which is the right way to ship automation over a reviewed one.

The result is that the strongest guarantee in the project is on the phase that needs no guarantees, since review is read only, and the weakest is on the phase that touches GitHub, which is why the gate is built there.

Publishing claims a slot and cleans up its own pending reviews

Posting is the irreversible act, so it gets the most careful engineering in the project. Everything goes through the GitHub CLI against the reviews endpoint, with a posting claim recorded first and self healing cleanup of leftover pending reviews, which the readme gives as the reason duplicate concurrent posts cannot happen.

That is a real failure mode rather than a hypothetical one. A review submitted and retried after a timeout can produce two comments, and the second one is the kind of error a human has to notice and delete manually. Claiming the post before making it, then sweeping anything left pending, means the retry path is safe.

The placement logic is also careful. A finding whose line reference lands on a line the pull request actually changed is posted as an inline review comment, because that is where a reviewer will read it. Everything else is collected into a summary section instead of being dropped, which is the second defensible choice here: a finding on an unchanged line is usually still a real problem, and losing it would make the review look shorter than the work.

The gate itself is per finding, not per run. Each finding carries a checkbox to post it as a pull request comment and an optional note, and that note is woven into the generated comment as an edit instruction rather than quoted, which keeps the audit trail of your objection while producing something that reads as a review.

Before any of that there is a preview, and it is explicitly a dry run that can be cached and regenerated. Language handling is part of the same step: findings written in whatever language you are working in are rewritten as professional English before they reach GitHub.

Commit and upload shows the diff before it runs git

The fix path is a persistent chat per pull request, and its permission model is stated in the first sentence: the agent edits the worktree but does not commit or push by default.

The action that changes that is named for what it does. Commit and upload first displays the diff and an editable conventional commit message, and only confirmation runs the add, commit and push sequence. So the agent's work is always inspectable before it becomes history, and the message is a field you can rewrite rather than one you accept.

Around that are the controls an autonomous loop needs. Stop and resume, expandable run logs, decision cards for the points where the agent needs an answer, an ultracode toggle that escalates the current turn to a higher reasoning effort, and a separate explicit toggle for dangerous commands.

The project assistant works the same way with one more permission. A session can start in the project directory or in an isolated feature worktree on a new branch, and real decision points are rendered as ask-user cards rather than being guessed at. Opening a pull request is treated as an explicit action that grants the agent permission to commit, push and run the pull request creation command for that turn only, and commit messages and the pull request title and body are always written in English regardless of the conversation language.

A global assistant at the bottom right inherits the project's provider and working directory when they are available and understands a small command set for ad-hoc work, which is how you troubleshoot a single repository without starting a review run.

The agent SDK is pinned exactly and the readme names a different package

The manifest pins the OpenAI package to a single version with no range at all:

code
"@openai/codex": "0.149.1"

Every other runtime dependency in the same list uses a caret range. One package out of fourteen is frozen at an exact version, and it is the one that wraps the provider whose behaviour changes most often. That is defensible for an agent SDK whose output format the application parses, and it is also the dependency that will need a manual bump first.

There is a second mismatch in the same area. The readme's stack description names the Codex SDK package, while the manifest depends on the Codex package without the SDK suffix. Two similarly named packages with different scope, one in the prose and one in the build, is exactly the kind of thing that sends someone looking for a dependency that is not there.

The rest of the stack is a Nuxt 4 application with a component library and Tailwind version 4, a SQLite database reached through an ORM, and an Electron wrapper that runs the server under Electron's own Node mode. The desktop build runs three preparation scripts in sequence, one for the SQLite native module, one for Codex and one that writes build information.

The test script is worth noting for a project this size. It is not a test runner: it is a shell loop over the test and contract files, invoking a TypeScript executor on each and exiting on the first failure. It is deterministic and dependency free, and it assumes a POSIX shell.

Turning the fast tier off takes two files, and the readme says so

The example environment file is the best documentation in the repository, because every line carries its explanation in three languages at once, Chinese, English and French, matching the three readmes.

Two settings decide where inference runs. One selects the provider globally, with the documented options being the locally logged in binary for one provider, which uses an existing subscription and therefore little extra API cost, and the other provider as the default that a project can still override. The second selects the default review model for the first provider, accepting the short aliases and a full model name, and the model for the second provider is left commented out so the tool's own default applies.

The service tier setting is the trap. The file explains that the project level Fast toggle overrides whatever is set globally, and then adds the part that is easy to miss: if the same option is also written in a configuration file in the user's home directory, that one has to be deleted or commented out as well.

So the cost-relevant setting for that provider lives in two places, and clearing one of them changes nothing. That is stated in the example file rather than discovered later, which is the right place for it.

The rest of the configuration surface follows the same discipline elsewhere in the project. Effort level, service tier and review methodology are configurable per project rather than globally, and each project picks one provider, with review, fix chat, recheck, skill generation and publish time rewriting all following that choice without mixing sessions or models. Skill activation is likewise guarded: a generated candidate is saved and diffed first, and activating it never overwrites the current skill.

Editorial conclusion

PR Cockpit is a careful design for the specific failure everyone has with agent assisted review, which is publishing a confident wrong comment under your own name. Its read only review agents, its per finding gate and its transactional publishing are the parts worth copying even if you use something else. Three practical cautions. You are running an unreleased nightly, so pin a known good build rather than tracking the latest. Nothing is signed, so clear the quarantine flag only on a machine where you would accept an unsigned binary anyway. And the human gate only applies to posting; the agent still edits worktrees, so keep the repository you point it at disposable.

Frequently asked questions

Are PR Cockpit releases versioned?

No. Builds are published as a rolling nightly pre-release rebuilt on every push to main, and the release list contains a single entry for that channel. There are no versioned tags to pin, and the manifest sits at 0.1.0 and is marked private.

Are the PR Cockpit desktop builds signed?

No. None of the packages are code signed, so macOS reports the app as damaged or unidentified, Windows shows the SmartScreen dialog, and the Linux AppImage has to be marked executable before it will start.

Can the PR Cockpit review agent modify files or push?

No. Review agents are read only, with tool level blocking for git writes, file edits, network access and dangerous commands, plus an operating contract and skill linting. Anything that writes runs in an isolated worktree instead.

What stops PR Cockpit from posting duplicate review comments?

Publishing goes through the GitHub CLI against the reviews endpoint with a posting claim recorded first and self healing cleanup of leftover pending reviews. Findings on changed lines become inline comments and the rest go into a summary section rather than being dropped.

Does the PR Cockpit fix agent commit and push on its own?

No. The agent edits the pull request worktree but does not commit or push by default. Commit and upload first shows the diff and an editable conventional commit message, and only confirmation runs the git sequence. Opening a pull request is a separate explicit per-turn permission.

What do I need before PR Cockpit can do anything?

A completed GitHub CLI authentication, since every GitHub read and write goes through it, plus a login for either Claude or Codex. Building from source additionally needs Node 22 or later and pnpm 9.

Official sources

  1. Issues
  2. License: MIT
  3. README
  4. Releases
  5. taovc/pr-cockpit on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/taovc-pr-cockpit.svg)](https://hysenlabs.com/projects/taovc-pr-cockpit)