open-computer-use: an open-source Computer Use MCP server for Codex, Claude Code and Gemini CLI
👾 Open Computer Use – Open-Source Alternative to Codex Computer Use
At a glance
- What is it?
- open-computer-use wraps Accessibility-based desktop control as an MCP service on macOS, Linux and Windows. Here is how it installs, what the CLI actually exposes, and where the macOS permission model becomes the real constraint.
- Who is it for?
- Adopt it if you already drive an MCP-capable agent and want desktop control without writing platform glue, and if you accept that on macOS everything depends on Accessibility and Screen Recording grants. Skip it if your target machine cannot grant those permissions, or if you need a documented rollback path: the README documents install commands but not uninstall ones.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Swift, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap open-computer-use fills between an agent and a desktop
Most agent tooling stops at the shell. An agent can read files, run commands and call HTTP APIs, but the moment a task lives inside a GUI application, the agent is blind. OpenAI's Codex Computer Use demonstrated that this does not require screen-scraping or a virtual machine: the README states the project was inspired by that work and that non-intrusive CUA can be built on top of Accessibility. open-computer-use takes that idea and packages it as an MCP service, so any MCP client can drive a desktop rather than only one vendor's agent.
The intended user is someone who already runs an agent that speaks MCP (Codex, Claude Code, Gemini CLI, opencode) and wants that agent to touch native applications. It is not a standalone automation tool. There is no scripting language of its own, no scheduler, no visual workflow builder. The agent supplies the reasoning; this project supplies the hands. The README frames it plainly: it is an open-source Computer Use service wrapped as MCP, and the repository carries topics for accessibility, gui-automation and cua alongside mcp and model-context-protocol.
Accessibility as the control surface, MCP as the transport
The mechanism is stated in the README and visible in the CLI surface. Rather than synthesising mouse events at fixed coordinates, the runtime reads application state through the platform accessibility layer. The CLI exposes this directly: open-computer-use call list_apps enumerates running applications, and open-computer-use call get_app_state --args '{"app":"TextEdit"}' returns the state of one application. That returned state carries an element_index, which is the handle the agent uses in subsequent actions.
That indexing detail explains one of the more interesting flags. The README notes that a sequence run in a single process lets element_index state be reused, and gives open-computer-use call --calls with a JSON array of tool invocations, plus --calls-file for reading the sequence from disk. Sequence runs sleep 1s between successful operations by default, and --sleep overrides that. The sleep is not cosmetic: GUI applications need time to repaint before the next accessibility read is meaningful. Running each action as a separate process would invalidate the indices, which is why the batched form exists.
Platform coverage is claimed for macOS, Linux and Windows, but the permission story is not uniform. On macOS the runtime requires macOS 14.0 or later, and the README instructs the user to run it once and grant Accessibility and Screen Recording. Windows and Linux do not need that step. So the same MCP server has two very different operational profiles depending on the host.
Installing open-computer-use and wiring it into Codex or Claude Code
The package is distributed through npm, with ocu as a short alias. Install it globally:
npm i -g open-computer-useOn macOS, run it once so the permission prompts appear, then grant Accessibility and Screen Recording in System Settings. The README states the macOS runtime requires macOS 14.0 or later. Windows and Linux skip this step entirely.
open-computer-use
# or
ocuRather than editing agent config by hand, the CLI writes it for you. For Codex it appends to ~/.codex/config.toml:
open-computer-use install-codex-mcpFor Claude Code the target is ~/.claude.json, and for Gemini CLI the current project's ./.gemini/settings.json (or the user config with --scope user):
open-computer-use install-claude-mcp
open-computer-use install-gemini-mcp
open-computer-use install-gemini-mcp --scope userIf you would rather register it manually in any MCP client, the README gives this JSON shape. The command is open-computer-use and the argument is mcp:
{
"mcpServers": {
"open-computer-use": {
"command": "open-computer-use",
"args": ["mcp"]
}
}
}A first real use is a single tool call that prints the MCP-style JSON result. This is the fastest way to confirm the server is alive and the permissions are correct, before an agent is involved at all:
open-computer-use call list_appsIf that returns an application list, the runtime is working. Then check permissions explicitly with open-computer-use doctor, which the README says opens onboarding only when something is missing. There is also a skill distributed separately through npx skills add iFurySt/open-codex-computer-use with -a codex or -a claude-code, and npx skills update open-computer-use -g -y to refresh an existing global install.
Where the macOS permission model becomes the failure mode
The sharpest limitation is not a bug, it is the platform contract. On macOS, both Accessibility and Screen Recording are user-granted, system-level permissions that cannot be scripted around. A headless CI runner, a fresh VM image, or a locked-down corporate laptop that blocks those grants will leave the tool installed and inert. The README is explicit that Windows and Linux do not need this step, which means the macOS path is the one with the manual gate, and it is also the platform with the 14.0 minimum.
The second limitation is the timing model. The default 1s sleep between successful operations in a sequence is a blunt instrument. It is a fixed wait, not a condition check, so a slow application can still be mid-repaint when the next accessibility read happens, and a fast one wastes a second per step. The --sleep flag lets you tune it, but tuning is on you, per application. Nothing in the README describes an adaptive or event-driven wait.
The third is scope. This is a Computer Use service, not a test framework. There are smoke and stress targets in the Makefile (make smoke, OPEN_COMPUTER_USE_STRESS_LOOPS=20 make stress) and agent smoke tests driven by scripts/run-agent-smoke-tests.mjs, but those validate the tool itself. If you want deterministic, assertion-based GUI tests with reporting, this is the wrong layer. It is also the wrong tool if you need the agent to run without a logged-in graphical session.
How it differs from Playwright-style browser automation
The obvious alternative for GUI automation is a browser driver such as Playwright or Selenium. The difference in approach is structural, not a matter of maturity. A browser driver controls one application (the browser) through a protocol that application implements for the purpose, and it addresses elements through the DOM. It is deterministic, fast, and headless-capable, and it has nothing to say about a native editor, a system dialog or a desktop application that has no web view.
open-computer-use addresses the operating system's accessibility tree instead, which is why it can reach TextEdit or any other native application, and why it inherits the accessibility tree's quirks: elements must be exposed by the application, and the element_index is a positional handle rather than a stable selector. That is the trade. You get breadth across native applications at the cost of determinism and speed, plus a fixed inter-step delay. The README also points to a sibling project, open-browser-use, for the browser case, which is a sensible split: use the browser driver when the target is a web page, and this when the target is the desktop.
Licence, upgrade path and what a version bump actually costs
The repository is MIT licensed, with a THIRD_PARTY_NOTICES.md file at the top level. MIT is permissive: it allows commercial use and modification, and it requires that the licence and copyright notice be preserved. That is the whole obligation in the licence text itself, but the project bundles third-party components, so the notices file is worth reading before redistribution. Nothing here is legal advice; if you ship this inside a product, read LICENSE and THIRD_PARTY_NOTICES.md in full.
Upgrade cost is low and mostly mechanical. The npm package is the distribution channel, so a global install moves forward with the package version, and the README documents npx skills update open-computer-use -g -y for refreshing the skill. The release cadence is visible: v0.3.5 on 2026-09-09, v0.3.4 on 2026-09-08, v0.3.3 on 2026-09-01, and the last push to the repository was on 2026-09-09. Those are pre-1.0 versions shipped days apart, which is a signal about interface stability rather than about quality: expect the CLI surface and tool arguments to keep moving, and pin the version if you depend on a specific tool name or argument shape.
The gap worth flagging is uninstall. The README documents install-codex-mcp, install-claude-mcp, install-gemini-mcp and install-opencode-mcp writing into ~/.codex/config.toml, ~/.claude.json, ./.gemini/settings.json and ~/.config/opencode/opencode.json respectively, but it does not document a removal command. If you install into a shared agent config, plan to edit that file by hand when you back out.
Editorial conclusion
Adopt it if you already drive an MCP-capable agent and want desktop control without writing platform glue, and if you accept that on macOS everything depends on Accessibility and Screen Recording grants. Skip it if your target machine cannot grant those permissions, or if you need a documented rollback path: the README documents install commands but not uninstall ones. Verify first that node and npm are available, that the machine meets macOS 14.0 or later, and that open-computer-use doctor reports no missing permission before you wire it into an agent config.
Frequently asked questions
How do I turn on computer use in Codex?
Install the package with npm i -g open-computer-use, then run open-computer-use install-codex-mcp, which writes the MCP entry into ~/.codex/config.toml. On macOS, run the binary once first and grant Accessibility and Screen Recording.
Can OpenClaw do computer use?
OpenClaw is not covered by the README, so there is nothing here to confirm about it. open-computer-use itself provides Computer Use through MCP for clients such as Codex, Claude Code and Gemini CLI.
Can I use Codex on my desktop with open-computer-use?
Yes. The README shows open-computer-use used as Computer Use in Codex App and Codex CLI, and there is a separate install-codex-plugin command described as mainly for Codex App. The runtime supports macOS, Linux and Windows.
Can Codex access my computer through open-computer-use?
It accesses applications through the platform accessibility layer rather than raw input injection. On macOS that requires the user to grant Accessibility and Screen Recording, and the runtime requires macOS 14.0 or later.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/ifuryst-open-codex-computer-use)