clawdcursor addresses the screen by id, then checks its own work
clawdcursor compiles whatever's on screen into one UI map — accessibility tree and OCR fused into stable, addressable elements, with a screenshot only when needed — then drives apps through reusable scripts, verifying every action and routing it through a single safety gate.
At a glance
- What is it?
- Clawd Cursor is a local MCP server that compiles the accessibility tree and OCR into one UI map of stable `el_NN` elements, so an agent acts on ids instead of pixels, optionally re-reads the screen to confirm the result, and passes every call through a single `safety.evaluate()` gate. The README positions it explicitly as a last-mile fallback, not a replacement for an API.
- Who is it for?
- clawdcursor earns a place when the work has no API and no CLI, especially legacy desktop apps and canvas surfaces, and its `expect` parameter plus the single gate are the right shape for supervised automation.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Elements get stable el_NN ids, so the agent stops sending coordinates
The central mechanic is a compile step. Clawd Cursor reads the accessibility tree and OCR output, fuses them into one set of elements scored by confidence, tags each with a stable id shaped like `el_NN`, and exposes that map to the agent. Actions are addressed by id.
That is the difference from the screenshot-and-vision approach, where the model emits a point and the tool replays it. Pixel targets drift as a window moves, a dialog opens or a theme reflows. An id survives those changes, so a recorded action can be replayed after the layout moves under it.
The README is explicit that coordinates are not gone, only demoted. They appear in the last-resort tier that takes live pixels off the current frame, reserved for canvas-only apps and for tasks that genuinely need spatial reasoning. There is no canvas accessibility tree to read, so for that class of app the expensive path is the only path.
The scope claim follows from the same design: native apps, the browser, or a canvas. What it does not claim is to be a better browser tool than a browser tool, since the readme calls Playwright MCP and browser-use adjacent and positions Clawd Cursor as the fallback for when no public surface exists.
The perception ladder only reaches pixels when a task forces it
Perception is tiered by cost, and the tiers are ordered cheapest first: the accessibility tree, described as free, then OCR, described as cheap, then the screenshot, described as expensive and the only tier that puts pixels into the model's context.
The claim about tokens follows directly. If an agent climbs the ladder only when it must, its context cost tracks the difficulty of the screen rather than the number of actions. A form with labelled fields is entirely reachable from the tree. A chart drawn on a canvas is not.
Note the README's phrasing that screenshot and vision are the same step. There is no separate vision call: pixels go into the model's context or they do not, which means the expensive tier is also the privacy-relevant one. Everything below it runs on the machine.
Vision-centric parsers such as OmniParser and UI-TARS are named in the comparison as things you build an agent around rather than install. Their tradeoff is that a screenshot is needed in the model for every observation, so cost is paid on every step, not only on the hard ones.
The expect parameter turns a reported success into a checked one
Most automation reports what it intended. Clawd Cursor accepts an `expect` argument on a consequential action and re-reads the live screen afterwards to check it, with a short settle window so an asynchronously updating UI has time to repaint before the check runs.
When the screen does not match, the tool reports a DEVIATION instead of a success. The framing matters: a completed task cannot be marked done on evidence that was already true before the agent acted. If your expected state existed a moment earlier, reading it proves nothing about whether the click worked.
This is opt in per call rather than global, which is the right shape for it. Reads do not need verification; the expensive, destructive operations do. Leaving `expect` off a consequential action returns you to trusting the tool's own report.
The comparison table credits this row to Clawd Cursor alone, marking the other four approaches, browser-use, Playwright MCP, OmniParser and UI-TARS, and computer-use, as having no deviation check at all. The authors describe it as one of three things that are genuinely rare in this space.
safety.evaluate() is the single door, and the gate itself needed hardening twice
Every call routes through one function before it touches the desktop: `safety.evaluate()`, which returns allow, confirm or block. The claim is architectural. It does not matter whether the request arrives from an editor over stdio, from an external agent over HTTP, or from the built-in loop, because the check is on the path rather than on the caller.
The on-screen guard is the human half of that. An on-screen banner reading desktop control in progress appears with a blinking red dot whenever an agent is driving, and double-clicking it stops the session. So a person at the machine has both notice and a kill switch, which matters because an agent controlling a real desktop is not a sandbox.
The release history is worth reading here. Of the three most recent tags, v1.5.10 on 2026-09-28 is described as keyboard gate hardening and v1.5.9 on 2026-07-03 as batch confirm-tier hardening, both marked as security work. A safety gate that needed two hardening passes in three months is a gate under real use, and it is also a signal that its earlier behaviour was not what the changelog said.
Third-party named peers are Windows-MCP and Terminator for desktop control, with computer-use noted as Claude only and sandboxed. Model independence is the other claim: any vendor, any model.
Install, consent and the macOS grant are three separate steps
The engine install and the permission grant are not the same command, and the readme separates them:
npm i -g clawdcursor
clawdcursor consent --acceptConsent is described as one-time desktop-control consent and marked required. A global install is also recommended over the zero-install path, where you swap `clawdcursor` for `npx -y clawdcursor` in any snippet, on the grounds that a global install is pinnable and inspectable. That is a reasonable default for a binary that can click anything on the machine, and pinning matters more here than for a typical CLI.
macOS adds a third step. `clawdcursor grant` is macOS only and is where you approve Accessibility and Screen Recording in the system dialogs. Those are the two permissions that let a process see the UI of other applications, so on that platform the consent record and the OS grant are separate systems and both have to be satisfied.
Windows and Linux are named as supported platforms, with the grant step scoped to macOS. The repository ships platform scripts under `scripts/mac/`, `scripts/linux/` and `scripts/*.ps1`, which suggests the setup path differs per OS rather than being one binary everywhere.
The postinstall chain cannot fail, which cuts both ways
The npm lifecycle script is written so that neither verification step can break an install:
"postinstall": "node scripts/verify-install.js || node -e \"process.exit(0)\" && node scripts/postinstall-native.js || node -e \"process.exit(0)\""Both branches end in `process.exit(0)`. If the verification script fails, the install continues; if the native helper fails, the install continues. The practical effect is that `npm i` reports success on a machine where the desktop integration is not actually usable, and the first sign of trouble arrives later as a tool call that cannot find elements.
That trade is common in tools that ship platform binaries, because a postinstall that hard fails on one OS poisons installs everywhere. The compensating step is a doctor command: `clawdcursor doctor` is what the environment example recommends to auto-detect and configure, and it is the right first thing to run after installing.
The build itself is plain TypeScript to JavaScript with a postbuild pass, and `clean` is a plain recursive remove of `dist`, so a rebuild is cheap.
The tarball ships Swift sources and an entitlements file, not an opaque blob
The `files` array in package.json includes `native/build.sh`, `native/Package.swift`, `native/Sources/` and `native/entitlements.plist` alongside `dist/` and the scripts that verify and post-install the native layer. The build step is plain enough to read:
"build": "tsc && node dist/postbuild.js",Shipping the Swift sources rather than only a compiled helper means the accessibility and screen recording side of the toolchain is inspectable before you let it run. `entitlements.plist` in particular is the kind of file worth reading, since it declares what the helper is permitted to do on macOS.
The repository has the supporting structure for that scrutiny: `tests/` with a separate `tsconfig.tests.json` and a vitest config, a `schema.snapshot.json` that pins the tool surface so a schema change shows up as a diff, `perf/` for measurement, and a `seed-registry/` directory. There is also a `SKILL.md`, a `server.json`, a `.claude-plugin/` entry and a `.env.example`.
Version numbers disagree with tags at the moment: package.json says 1.5.11, and the newest published tag is v1.5.10 from 2026-09-28, with the last push on 2026-10-01.
Editorial conclusion
clawdcursor earns a place when the work has no API and no CLI, especially legacy desktop apps and canvas surfaces, and its `expect` parameter plus the single gate are the right shape for supervised automation. Check three things first: that you ran `clawdcursor consent --accept` deliberately, whether the macOS Accessibility and Screen Recording grants were approved, and which tier a given screen forces you into, since a canvas-only app means the screenshot tier and a real per call token cost. Note that package.json sits at 1.5.11 while the newest tag is v1.5.10, and that two of the last three releases were security hardening on the gate itself.
Frequently asked questions
How does clawdcursor address elements on the screen?
It fuses the accessibility tree and OCR into a confidence-scored map and tags each element with a stable el_NN id, then acts on the id rather than on pixel coordinates. Coordinates appear only in the last-resort screenshot tier for canvas-only apps or spatial tasks.
What does the expect parameter do in clawdcursor?
Pass expect on a consequential action and Clawd Cursor re-reads the live screen afterwards, with a short settle window for async UIs. If the screen does not match it reports a DEVIATION rather than reporting success.
Do I have to approve permissions on macOS to use clawdcursor?
Yes. Besides the one-time clawdcursor consent --accept step, macOS needs clawdcursor grant to approve Accessibility and Screen Recording in the system dialogs. The grant step is scoped to macOS.
Does clawdcursor need a model API key?
The MCP server itself is local and model independent, so an agent that brings its own model needs no key here. The bundled loop does take one: AI_API_KEY plus an optional AI_PROVIDER of anthropic, openai, ollama or kimi, auto-detected from the key format when unset.
How does clawdcursor stop a runaway agent on my desktop?
Every call routes through a single safety.evaluate() chokepoint that returns allow, confirm or block, and the caller cannot bypass it. An on-screen desktop control in progress banner appears while an agent drives, and double-clicking it stops the session.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/amrdab-clawdcursor)