# snip: an Electron diagram and screenshot renderer whose MCP server only starts from a source checkout

> A menu bar app plus a CLI that lets a coding agent render Mermaid diagrams and HTML previews for human review, with annotations fed back as a structured verdict. It ships through Homebrew, a disk image, an AppImage and a deb, and two of its capabilities only work if you cloned the repository.

**rixinhahaha/snip** — The visual communication layer between humans and AI agents. Capture, annotate, render diagrams, and organize with AI — powered by Electron and Ollama. macOS & Linux.   

- Repository: https://github.com/rixinhahaha/snip
- Website: https://snipit.dev
- Stars: 326 · Forks: 31
- Language: JavaScript
- License: MIT
- Published: 2026-09-15 · Updated: 2026-09-15 · Language: en
- Canonical page: https://hysenlabs.com/projects/rixinhahaha-snip

## The MCP entry point is a file path into a clone, launched with the system node

The configuration block for agents without shell access is the whole interface:

```json
{
  "mcpServers": {
    "snip": {
      "command": "node",
      "args": ["/path/to/snip/src/mcp/server.js"]
    }
  }
}
```

Three things follow. The command is the system node, not the one inside the application. The argument is a path into a source checkout, with no build step, no resolved install location and no version. And nothing in that block tells the server where the application lives, so it is reaching for the same user data directory the packaged app uses.

That last part is the real constraint. The four documented install channels are a Homebrew cask, a disk image, an AppImage and a deb. None of them puts a JavaScript file where this configuration can point at it.

There is a second wrinkle underneath. The manifest declares the package entry point as the Electron main process file and the console script as a file inside the CLI directory, so the package is shaped like an application rather than a library. It also declares a native addon binding and depends on a native interface package at runtime, and there is a build configuration file at the repository root to go with them.

An MCP server started by the system node is not running under Electron, so any native module it touches has to be compiled against the system node rather than the bundled one. The repository has a rebuild script for exactly that, and an install hook that runs the rebuild only on macOS. On any other platform the hook is a no-op.

The consequence is that the documented agentless path is a developer path, and it is not marked as one.

## The macOS build targets one architecture and signs ad hoc

The build script is one line and it is narrow. It rebuilds the native modules, then sets automatic code signing identity discovery to false, then builds a macOS arm64 target, then runs a separate ad hoc signing step.

The signing step loops over the built application bundles and, for each one that is not already validly signed, forces a deep ad hoc signature. An ad hoc signature identifies no developer, carries no certificate chain and is not notarised.

That sits directly against the sponsor section, which says the Apple Developer Program membership is required for code signing and notarisation so that macOS does not block the application. The build script does not do the thing the sponsor money pays for.

There is no x86_64 macOS build script, so an Intel Mac has no build target in the manifest even though the disk image note calls out Apple Silicon as the one to download.

Linux has its own script and it is x86_64 only, matching the artifact names in the install section: one AppImage for x86_64 and one deb for amd64.

So the two supported platforms have disjoint architectures. macOS is Apple silicon, Linux is x86_64, and neither has the other. That is unusual enough to check against your own machine before you plan an install.

## Linux is described three ways: Wayland only, x86_64 only, and any distro

The development requirements line says macOS 14 or newer, or Linux on Wayland, with Node 18 or newer. Wayland is the whole Linux story; there is no mention of an X11 session anywhere.

The install section then offers an AppImage described as portable and any distro, and a deb for Ubuntu and Debian. Those descriptions are about packaging, not about the display server, and a Wayland-only application distributed as an any-distro AppImage is a mismatch worth noticing rather than a contradiction.

The same install section is precise about architecture. Both Linux artifacts are named with an x86_64 or amd64 suffix and there is no arm64 Linux build anywhere in the manifest scripts.

The keyboard shortcuts section adds one more platform fact. On Linux the command key becomes the control key, which covers every shortcut in the table, and the documentation states that as a single rule rather than repeating it per row.

The result is a Linux story that requires a Wayland session on an x86_64 machine, which rules out most ARM desktops, most X11 sessions, and a Raspberry Pi, while the macOS story requires an Apple silicon machine.

## The segmentation model is a dev dependency fetched by a script, not shipped

The screenshot editor has eight annotation tools, and the last one is an AI segment tool that uses a segmentation model. The tech stack names it as a small model in the ONNX format, and the library that runs it is a Transformers.js package.

That library is in the development dependencies, not the runtime ones, and the models are fetched by a dedicated download script. There is a second script for downloading a node runtime, which is presumably for the packaging pipeline.

So on a packaged install the segment tool has neither its runner nor its weights, and the only documented way to get them is a script in a repository you have not cloned.

This is the same shape as the MCP problem in the first section, and it points at the same root cause. The application is distributed as a finished binary while two of its capabilities assume a working directory.

The contrast with the rest of the stack is instructive. The other AI feature, organising and searching the screenshot library, runs against a local model server that the user installs separately, and the README says so with a download link. That one is honestly external. The segmentation model is internal and therefore ought to have been bundled or downloaded by the application itself.

One more command has no visible engine. There is a transcription command described as extracting text from an image with optical character recognition, and no OCR library appears in the runtime dependency list, which holds a canvas library, a diagram renderer, a local model client, an image encoder and a decoder, two PNG paths, a file watcher and the updater.

## Nine agent tools for ten commands, and one tool with no command behind it

The command table has ten rows. Setup configures the agent. Render takes one of two format flags and reads from standard input. Open, search, transcribe, list, get, organise and categories are the rest.

The agent tool list has nine names. Diagram rendering, opening in the app, searching, listing, getting metadata, transcription, organising, categories, and installation of an extension.

The mapping is not one to one in either direction. The two render formats collapse into a single diagram tool, so an agent that wants an HTML preview has no documented argument for it. And the extension installation tool has no command-line counterpart at all, so whatever it does can only be reached through the agent path.

The names are also inconsistent in a small way that matters for anyone writing a tool selection policy. The command line says organise and the tool says organise a screenshot; the command line says categories and the tool says get categories. Both are fine, and both are exactly the sort of drift that makes an agent pick the wrong one.

The return value is where the two surfaces agree, and it is the most useful thing in the file. The agent gets a small object with a status that is either approved or changes requested, whether the image was edited, a path, and text. That is a review loop with a machine-readable verdict, which is the whole point of the project.

## Review commands block on a human, so an unattended agent call waits

One sentence states the behaviour plainly: the review commands block until the user finishes and return structured JSON.

That is the correct design for the intended use, where a person is looking at the screen and approving or annotating. It also means the command has no timeout, no non-interactive mode, and no documented escape.

For an agent that runs unattended, a blocking review call is a hang. The only escape in the file is an environment override for local quality assurance that forces mock practice and deterministic sessions, and the surrounding instructions tell you to use it on a trusted local machine only.

So there are three ways the review can end: a person approves, a person requests changes, or you switch the application into a scripted mode for testing. Nothing documents a way to answer the review automatically in normal use, which is a reasonable thing to leave out of a screenshot tool and a real limitation for a CI pipeline.

The annotation surface is what makes the verdict useful. You can annotate spatially, type text feedback, or approve without touching anything, and the editor supports eight tools plus a three-level undo where annotations are undone before a crop and a crop before a segment cutout.

## Setup rewrites the agent's rules, its skills and its permissions

The setup command is described as the whole integration. It configures the coding agent to use visual output automatically, and the README enumerates what it touches: rules, a diagram skill, and permissions.

So a single command modifies three separate kinds of configuration in another tool. Rules are agent instructions, the skill is a slash command, and permissions are the list of things the agent may do without asking. Those are exactly the three files a careful user would want to review before agreeing to.

There is no dry run, no list of what will change, no backup, and no uninstall command in the ten row table. Undoing it means editing the agent's configuration by hand or reinstalling it.

The advertised scope is wider than one agent. The README says built for one coding CLI, and also working with three others and anything that can run a shell command, and the setup command's own description lists the same four. That is a reasonable claim for a shell-driven integration and it means the configuration being rewritten is whatever those four tools read.

The slash command is the discoverable part. Typing it in any conversation visualises whatever is being discussed, and the setup step is what makes the automatic behaviour fire when you ask about architecture, data models or flows.

## Five months of silence, and a tag cut sixteen hours after the last commit

The version is 1.3.14. The three most recent tags are patch releases on consecutive days in April and then one in early May: 1.3.12, 1.3.13 and 1.3.14.

The last commit is 2026-05-07 in the morning and the 1.3.14 tag is that evening, about sixteen hours later. Nothing has been committed since, which is close to five months. That is short of the half year at which a project would need to be described by its last commit date rather than as moving, but it is a long silence for an application that calls itself built by two developers and asks for sponsorship to cover ongoing work.

The repository root is unusually broad for an application. Alongside the source, tests and build configuration there is a documentation directory with five named documents, a marketing directory, a site directory, a shell cleanup script, a directory listing manifest for an agent tool directory, and an agent instruction file.

Two details in the manifest and the readme are worth noting for anyone automating a build. The install hook that compiles native dependencies is gated to macOS, and the build hook that compiles them is gated the same way, so neither fires on Linux.

The readme also opens with an empty link element carrying campaign tracking parameters and no content, and contains two empty centred paragraphs where screenshots would go. The screenshot table at the end of the demo section has a header row and no images. The documentation it links is where the detail lives, not the readme.

## Conclusion

snip suits someone who already lives in a terminal with a coding agent and wants a diagram in the review loop instead of a paragraph, since the verdict comes back as structured JSON and the same surfaces exist over MCP. Two things to know before you rely on it. The agentless path needs a clone, because the MCP entry point is a raw file path into the source tree, and the segmentation model is fetched by a development script rather than shipped. On Linux it also needs Wayland and an x86_64 machine, which is a narrower claim than the package names suggest.

## FAQ

### What does snip do for an AI coding agent?

It renders Mermaid diagrams and HTML previews from standard input and opens them for human review, then returns a structured object with a status of approved or changes requested, whether the image was edited, a path, and text. Annotations drawn on screen are fed back so the agent can iterate. It also captures screenshots, annotates them, and searches the library.

### How do I install snip?

On macOS through a Homebrew cask or a disk image for Apple silicon. On Linux through an AppImage for x86_64 or a deb for amd64. Development needs macOS 14 or newer, or Linux on Wayland, with Node 18 or newer. There is no arm64 Linux build and no x86_64 macOS build script.

### Does the snip MCP server work from an installed app?

The documented configuration runs the system node against a path inside a source checkout, with no build step and no resolved install location. The four install channels do not put a JavaScript file where that path can point. An agent with shell access uses the CLI instead, which the README recommends as the primary integration.

### What does `snip setup` change in my coding agent?

Three things, per the README: the agent's rules, a diagram skill exposed as a slash command, and its permissions. There is no dry run, no file listing and no uninstall command in the documented CLI table, so reversing it means editing the agent's configuration by hand.

### Does snip send screenshots or images to a cloud service?

The README states all AI runs locally with no cloud APIs needed for core features. Screenshot organisation and search use a local vision model through a separate local model server you install yourself. The segmentation model behind the AI segment tool is a different path: its runner is a development dependency and its weights are fetched by a download script in the repository.

## Sources

- [License: MIT](https://github.com/rixinhahaha/snip/blob/main/LICENSE)
- [Project website](https://snipit.dev)
- [README](https://github.com/rixinhahaha/snip/blob/main/README.md)
- [Releases](https://github.com/rixinhahaha/snip/releases)
- [rixinhahaha/snip on GitHub](https://github.com/rixinhahaha/snip)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/rixinhahaha-snip
