# nuphus-mcp keeps native window handles on the far side of the boundary

> nuphus-mcp is a Rust desktop automation server speaking JSON-RPC over stdio, exposing 45 tools for screen, window, input, and Chrome control to any MCP client. Its most interesting design choice is the semantic desktop surface, where seven tools drive native UI through the accessibility tree and never let a window handle cross the protocol boundary.

**mrpulor-gh/nuphus-mcp** — Desktop automation MCP server — computer use for any AI agent: control screen, windows, mouse/keyboard, and Chrome via Model Context Protocol (stdio). Not a DSH plugin — DSH users install dsh-nuphus-mcp.

- Repository: https://github.com/mrpulor-gh/nuphus-mcp
- Website: https://github.com/mrpulor-gh/nuphus-mcp
- Stars: 316 · Forks: 38
- Language: Rust
- License: MIT
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/mrpulor-gh-nuphus-mcp

## Seven semantic tools keep coordinates off the protocol boundary

The most distinctive part of the tool surface is a set of seven tools that do not work in pixels at all.

They drive native user interface elements through the platform accessibility tree: Windows UI Automation on Windows and the accessibility API on macOS. The tool names describe the contract in order, listing targets, binding a target, observing semantically, picking a candidate, executing an action, performing a semantic action, and verifying state afterwards.

The sequence they impose is observe, pick a candidate, execute, verify. You cannot jump straight to an action, because the execution step takes a candidate that observation produced.

The boundary rule is what makes it more than a convenience. Native window handles, coordinates, and platform locators never cross the MCP boundary. A client sees opaque identifiers and stable semantic locators, and nothing about the windowing system leaks through.

That has a practical consequence. A client that has cached a coordinate from a screenshot cannot feed it back into these tools, and a session survives window moves or layout changes in a way that a purely pixel-based surface does not.

It is also why the tool count is described as two numbers rather than one: the semantic surface is a parallel path to the coordinate one, not a replacement for it.

## Vision is bring-your-own and fails loudly when it is unset

Exactly one tool needs a credential, and the documentation is precise about how to supply it.

The vision tool sends a screenshot to a model you choose. It speaks two protocols. The default is an OpenAI-compatible chat completions endpoint, which is what makes local and Chinese providers usable without changes. The alternative is the Anthropic native messages API, selected either by pointing the base URL variable at that host, in which case the protocol is inferred from the host, or by setting the provider variable explicitly.

Four environment variables back it. The key is required, the model identifier is required, the base URL is optional with a default pointing at the OpenAI endpoint, and the provider is optional with a default of automatic inference.

What matters is the failure mode. Nothing is required unless you call the tool, and when the tool is not configured it returns a clear error instead of failing silently. A server that quietly returned empty vision output would be worse than one that refuses.

The listed model examples span a hosted frontier model, a hosted vision model, and a local model, which is the same range the OpenAI-compatible default is chosen to cover.

## Local OCR downloads its models on the first run

The perception tool is the counterpart to vision, and it needs nothing from you.

It runs optical character recognition locally with a local OCR engine, and its models download automatically on first run. Icon detection is available as an optional addition using a separate detector model.

The pairing is the point. The vision tool answers what is in the frame semantically; the perception tool returns precise coordinates. Used together, an agent gets semantic understanding and pixel-accurate targets from the same screenshot.

That flow is described as battle-tested and inherited from the Nuphus desktop application, the parent project, rather than invented for the server. The parent is described as a local-first agent with real desktop execution and phone-as-second-screen sync, and this server exists to expose the same automation to clients that are not that app.

The first-run download is the one piece of setup friction in an otherwise keyless path, and it is worth knowing before you run it on a machine with no outbound access.

OCR also being local is what lets the server claim that desktop and browser automation need no API key at all.

## Windows is the only platform with full desktop control

The support table is short and the asymmetry is in the desktop column.

Windows gets full browser tools and full desktop tools through the Windows API. macOS gets full browser tools, and desktop support limited to the semantic accessibility surface plus mouse and keyboard, which requires the platform accessibility permission to be granted. Linux gets browser tools as available rather than full, and desktop support marked partial because window and input capabilities are limited.

The prerequisites section says outright that Windows is recommended for full desktop control, so the table is not underselling itself.

Browser tools have their own requirement: Chrome or Edge must be installed. The server auto-detects an installed browser, and if it finds none, the browser tool family returns a clear error rather than hanging.

That error path matters for the stdio model. A client launching the server has no way to display a dialog, so a missing browser has to surface as a tool result rather than as a startup failure, and that is what the documentation describes.

One more platform detail is specified precisely. Execution feedback is non-intrusive everywhere: Windows shows a compact status card anchored to the bottom right of the work area, clear of the taskbar and tray, reporting start, done, and fail in real time, while macOS and Linux post a system notification on completion, which can be turned off with an environment variable.

## The workspace is three crates and the transport is one line of JSON

The repository is a Cargo workspace with three members, and the split is by responsibility rather than by feature.

One crate is the server itself and is the product. One is the browser automation core, built on a Chromium protocol client. One is the desktop control core, and it is marked as vendored, built on screen capture and the Windows API with explicitly no dependency on the Tauri framework.

That last note is a design boundary. The parent application uses a desktop framework, and the server deliberately does not, so it can be a single binary with no GUI toolkit behind it.

The transport is described in the same breath: the process reads single-line JSON from standard input and writes responses to standard output, with no HTTP server and no daemon. The architecture diagram shows a client on the left, the server in the middle, and the two core crates pointing at screen, window, mouse, keyboard, and Chrome respectively.

The workspace manifest carries license, repository, homepage, and documentation metadata, and a comment explains that it exists as a template so future crates can inherit it. Every URL points at the repository rather than at a hosted documentation site.

## An npm package ships beside a Rust build requirement

Two packaging stories sit side by side.

There is an npm directory at the repository root, and the badges at the top of the README point at a scoped npm package. So the server is distributed to the JavaScript ecosystem as an installable package.

The prerequisites, however, say a stable Rust toolchain and tell you to build from source with Cargo. That is the requirement for anyone building the workspace directly rather than installing the published package.

Both are consistent, but they mean the install story differs by audience. A client written in JavaScript installs an npm package that carries or fetches the binary, while someone reading the repository builds the binary themselves.

The repository also carries a Cargo configuration directory, a lock file, a changelog, a contributing guide, a security policy, and two parallel tool references in English and Chinese. The dual documentation is not an afterthought: there is a mirror on a Chinese forking service for faster access inside that region, and the Chinese docs are served by default there.

For anyone tracking changes, the changelog plus three releases inside a fortnight is the fastest way to see how fast the surface is moving.

## It is a plain stdio server and not a plugin

One callout exists to prevent a specific mistake, and it is worth repeating.

If you use a particular harness framework, this server is not the thing to install. It is a plain stdio MCP server and not a plugin for that harness's plugin system, and a separate dedicated plugin package exists for those users.

The reason is structural rather than philosophical. A plugin runs inside a host process and gets that host's lifecycle, configuration, and logging. A stdio server is spawned by the client and communicates over pipes. Installing the server as a plugin would give it a lifecycle it does not expect.

The other structural choice is safety. Destructive tools are annotated as the specification requires, there is an optional strict-confirm mode, and paths are validated for screenshots, uploads, and file drags.

Path validation is the one that matters most for an agent with screen and keyboard access. An agent that can write the clipboard and drag files can do damage with no network involved, and validating the path before the operation is cheaper than reviewing the damage afterwards.

One more line in the platform section is worth noting as a design commitment: window activation is never used as a visibility fallback. The agent will not bring a window forward to see it, which keeps actions invisible to the person using the machine.

## Conclusion

nuphus-mcp suits anyone wiring computer use into an MCP client that already has a screen model, since local OCR needs no key and the semantic tools return stable identifiers instead of brittle coordinates. It does not suit a Linux desktop where window and input support is only partial, or anyone expecting a hosted service, since there is no daemon and no network listener. Before you adopt it, decide which vision provider you will supply, check that your platform is in the support table, and read the tool annotations if you plan to expose it to an agent that clicks things.

## FAQ

### Is MCP just a JSON?

For this server, effectively. It speaks JSON-RPC 2.0 over stdio: the process reads single-line JSON from standard input and writes responses to standard output, with no HTTP server and no daemon, just one binary. Any MCP client can connect and start controlling the screen.

### What is MCP and why is it used?

In this project's terms it is the transport that lets any MCP client, including Claude Desktop, Cursor, VS Code, Copilot, or the Nuphus agent itself, immediately control the screen, windows, keyboard and mouse, and drive Chrome. The same desktop and browser automation the Nuphus app uses is exposed to clients that are not that app.

### How many tools does nuphus-mcp expose?

Forty-five in total, split as 22 desktop and 23 browser, with the full reference in TOOLS.md and a Chinese counterpart. Browser tools cover navigation, accessibility-tree snapshots with references, click, type, script execution, scroll, extract, screenshot, evaluate, history, waiting, cookies, upload, tabs, and downloads.

### Which platforms does nuphus-mcp support for desktop control?

Full on Windows through the Windows API. On macOS it is semantic plus mouse and keyboard through the accessibility API, which requires the platform Accessibility permission. On Linux it is partial, with window and input capabilities limited. Browser tools are full on Windows and macOS and available on Linux.

### Do I need an API key to use nuphus-mcp?

Not for desktop or browser automation, and the built-in OCR is local. Only the vision tool needs one: the vision key and model are required, the base URL and provider are optional with defaults, and when the tool is unconfigured it returns a clear error rather than failing quietly.

### What is the Execution HUD in nuphus-mcp?

Non-intrusive feedback on every platform. Windows shows a compact status card anchored to the bottom right of the work area, clear of the taskbar and tray, reporting start, done, and fail in real time. macOS and Linux post a system notification on completion, which can be disabled with an environment variable.

## Sources

- [License: MIT](https://github.com/mrpulor-gh/nuphus-mcp/blob/master/LICENSE)
- [mrpulor-gh/nuphus-mcp on GitHub](https://github.com/mrpulor-gh/nuphus-mcp)
- [Project website](https://github.com/mrpulor-gh/nuphus-mcp)
- [README](https://github.com/mrpulor-gh/nuphus-mcp/blob/master/README.md)
- [Releases](https://github.com/mrpulor-gh/nuphus-mcp/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/mrpulor-gh-nuphus-mcp
