lahfir/agent-desktop: giving agents a real accessibility tree to work from
Agent Desktop gives any agent reliable computer use on the desktop. Built with Rust, it sees any app's real UI structure through OS accessibility trees and operates it — refs stay stable and actions stay safe to retry, instead of guessing from pixels.
At a glance
- What is it?
- A Rust CLI that observes macOS app UI through OS accessibility trees and acts on stable element refs, rather than guessing coordinates from screenshots. Here is how it installs, where progressive skeleton traversal fits, and when it is the wrong tool.
- Who is it for?
- Adopt agent-desktop if your agent targets macOS desktop apps and you want element refs instead of pixel coordinates, and you can grant Accessibility permission in the environment where the agent runs. Skip it if you need Linux or Windows parity today, or if your target is a browser page rather than a native app, unless you are prepared to drive Chromium through the CDP path with launch --cdp.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 3 days ago.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem: pixel guessing versus a real UI tree
Screenshot-driven computer use has a structural weakness. The agent sees a rendering of the interface, not the interface itself, so the coordinates it learns are tied to a window size, a scroll position and a theme. Move the window and the coordinates are stale. The README frames agent-desktop's answer in one line: it "sees any app's real UI structure through OS accessibility trees and operates it." That is the whole pitch, and it is a meaningful difference rather than a marketing one. An accessibility tree is the same structure screen readers consume, so it already exists in any app that supports assistive technology. The agent reads element roles, labels and hierarchy instead of inferring them from pixels.
The audience is narrow and specific. This is for people building computer-use agents that have to operate real desktop applications, particularly dense ones. The README names Finder, Safari, System Settings, Xcode and Slack as targets. If your automation lives entirely inside a browser tab, or inside an API that already exists, you are not the intended user.
Snapshots, qualified refs and why retries are safe
The mechanism has two parts. First, a snapshot walks the accessibility tree and returns a set of interactive elements, each with a ref and a snapshot_id. Refs look like `@s8f3k2p9:e1` and `@s8f3k2p9:e2`, so the ref carries the snapshot it came from. Second, actions address elements by that ref rather than by screen position. The README describes these as "qualified element references." The consequence is that an action is retryable: if the click fails, the same ref still points at the same element, provided the snapshot is still valid.
Interactions are headless by default. The README states that ref actions use accessibility APIs and "block silent focus, cursor, keyboard, or pasteboard side effects." In practice that means clicking a button through a ref does not hijack the user's cursor or overwrite their clipboard, which matters if the agent runs on a machine someone is also using.
Output is structured JSON with error codes and recovery hints. For an agent loop, that is more useful than a screenshot diff, because the failure carries a machine-readable reason instead of an image the model has to interpret.
Progressive skeleton traversal for dense apps
A full accessibility snapshot of a large app is expensive in tokens. The README includes a Slack comparison image captioned with 30,743 tokens for a regular snapshot against 383 for a skeleton overview. Those are the project's own figures from its documentation, not an independent measurement, and the gap will vary by app and by how deep the tree goes.
The technique is a shallow overview followed by targeted drill-down. `snapshot --skeleton` returns a depth-3 map where truncated containers report a children_count, and each truncated branch exposes a safe drill ref. You then re-snapshot with `--root` pointing at that ref and the original snapshot id. The README claims 78 to 96 percent token reduction on dense apps with this approach. Treat that as a range the maintainers publish, and check it against your own target app before designing around it.
The trade-off is round trips. Four commands instead of one means four process spawns, and the agent has to decide which branch to drill into. On a simple app, a full snapshot is cheaper in both tokens and latency. The README says as much: "For simple apps, a full snapshot is fine." Choosing the wrong mode is the most likely way to make this tool feel slow.
Installing agent-desktop and taking a first snapshot
The recommended install is npm, which downloads a prebuilt binary. The README notes this works without a global install too, via npx.
npm install -g agent-desktopnpx agent-desktop snapshot --app Finder -iBuilding from source requires Rust 1.89 or newer and macOS 13.0 or newer, per the README. The workspace pins `rust-version = "1.89"`, so an older toolchain will refuse the build.
git clone https://github.com/lahfir/agent-desktop
cd agent-desktop
cargo build --release
cp target/release/agent-desktop /usr/local/bin/Before anything works, macOS needs Accessibility permission. Screenshots additionally need Screen Recording, and the Notification Center opener needs Automation permission for System Events. The README states that plain permission checks never prompt, and that missing permissions should be requested in a bounded isolated helper:
agent-desktop permissions --requestPermission output is structured, with each capability as an explicit object. Automation can report `granted`, `denied` or `unknown`, where `unknown` means macOS would need to prompt or System Events could not be probed without prompting.
{
"accessibility": { "state": "granted" },
"screen_recording": { "state": "denied", "suggestion": "Grant Screen Recording permission" },
"automation": { "state": "unknown" }
}Once permission is granted, the basic loop is observe, act, re-observe. The README's simple-app example takes a snapshot of Finder, clicks a button by ref, types into a field, sends a keyboard shortcut, then re-observes.
agent-desktop snapshot --app Finder -i
agent-desktop click @e3 --snapshot s8f3k2p9
agent-desktop type @e5 --snapshot s8f3k2p9 "quarterly report"
agent-desktop press cmd+s
agent-desktop snapshot -iThe snapshot_id in those commands is the one returned by the first call. If you reuse a stale id, the refs no longer describe the current UI.
The command surface, and the four names that fail closed
The README counts 58 command names, of which 54 are operational. The categories are observation, interaction, keyboard, mouse, notifications, clipboard, window management, session lifecycle, trace read and export, plus a bundled `skills` doc loader. That is a wide surface for a CLI, and it reflects how much of desktop automation is not clicking: reading the clipboard, managing windows, firing notifications.
The four remaining names are reserved for held input and are documented as failing closed in the stateless CLI. Held input means a key or button that stays down across commands, which requires state that a process-per-command model cannot carry. The README says those names are reserved for a stateful daemon. So if your workflow depends on holding a modifier across separate invocations, the stateless CLI is the wrong entry point, and the README does not document an alternative daemon you can start today.
FFI bindings versus forking the CLI per call
Every GitHub Release ships a prebuilt C-ABI cdylib named `libagent_desktop_ffi` for macOS, Linux and Windows, alongside the CLI tarballs. The point is to avoid fork-exec per command. You `dlopen` the library and call the functions declared in `agent_desktop.h` in-process. The README lists Python, Swift, Go, Ruby, Node and C as consumers.
The Python example checks the ABI major before any call, via `lib.ad_init(4)`, then creates an adapter, runs a snapshot, parses a qualified ref, executes by ref, and destroys the adapter. Full details on entrypoints, ownership, threading, error handling, build and link, release archives and verification live in `skills/agent-desktop-ffi/`. That directory is the place to look before writing binding code; the README is not a substitute for it.
Note the platform mismatch worth watching. The FFI artifacts cover macOS, Linux and Windows, but the workspace members are `crates/core`, `crates/macos`, `crates/windows`, `crates/linux`, `crates/ffi` and `src`. The README's installation section and permission discussion are macOS-only, and the feature list names macOS-specific permissions. How complete the Windows and Linux backends are is not documented in the README.
CDP interop, and where this tool stops
For Chromium-based apps, `launch --cdp` opens a verified DevTools port so any framework that speaks CDP can drive the web contents, while native menus, dialogs and windows stay on the accessibility path. The README names Playwright, Puppeteer, `chrome-remote-interface` and agent-browser as examples. This is a sensible split: web content is already well served by CDP tooling, and duplicating it through accessibility would be worse.
That split also marks the boundary. If your automation target is a web page, use CDP directly and skip this tool; agent-desktop's value is in the native chrome around the page. If your target is a native app with a poor or absent accessibility tree, this approach degrades toward the pixel guessing it was meant to replace, and the README does not describe a fallback for that case.
The permission model is the other hard constraint. Accessibility permission has to be granted on the machine where the agent runs, and the README is explicit that plain checks never prompt. In a CI container or a headless server without a logged-in macOS session, there is no tree to read. This is a desktop tool, not a server tool.
Licence, maintenance and upgrade cost
The project is Apache-2.0, and the workspace declares `license = "Apache-2.0"` in Cargo.toml. Apache-2.0 includes an explicit patent grant and permits commercial use and modification; it also requires that you preserve notices and state changes. That is a permissive licence, but if you redistribute a binary that links the cdylib, read the licence text and your own legal counsel rather than treating this paragraph as advice.
The last push to the default branch was on 2026-09-08, and the most recent release listed is v0.8.5 from 2026-09-06, preceded by v0.8.4 and v0.8.3 in the weeks before. Release cadence has been roughly weekly. The workspace version in Cargo.toml is 0.9.1 with a release-please marker, which indicates automated versioning driven by conventional commits. For an adopter, that means version numbers move and the changelog is the place to read before upgrading. Note the version skew between the workspace manifest and the latest tagged release; do not assume a tag exists for every version string you see in the source tree.
Upgrade cost is mostly in the refs contract. Refs are qualified by snapshot_id, so a snapshot taken before an upgrade and replayed after one is not guaranteed to resolve. Pin the version you build against and re-snapshot after upgrading rather than replaying stored refs. The four held-input names that fail closed are another thing to re-check on upgrade, since a future daemon could change their behaviour.
Editorial conclusion
Adopt agent-desktop if your agent targets macOS desktop apps and you want element refs instead of pixel coordinates, and you can grant Accessibility permission in the environment where the agent runs. Skip it if you need Linux or Windows parity today, or if your target is a browser page rather than a native app, unless you are prepared to drive Chromium through the CDP path with launch --cdp. Before committing, verify three things on your own machine: that agent-desktop permissions reports accessibility as granted, that a full snapshot of your densest target app returns refs you can click, and that the FFI route you intend to use matches the ABI major you check with ad_init.
Frequently asked questions
What is agent-desktop and what does it do?
It is a Rust CLI that gives agents computer use on the desktop by reading an app's real UI structure through OS accessibility trees and operating it through element refs. The README describes the workflow as observe, decide, act, with structured JSON output and error codes.
How do I install agent-desktop on macOS?
The recommended route is npm install -g agent-desktop, which downloads a prebuilt binary, or npx agent-desktop for a one-off run. Building from source needs Rust 1.89 or newer and macOS 13.0 or newer, per the README.
What permissions does agent-desktop need?
macOS requires Accessibility permission, screenshots also require Screen Recording, and the Notification Center opener requires Automation permission for System Events. The README says plain permission checks never prompt, and that agent-desktop permissions --request asks for missing ones in an isolated helper.
How is agent-desktop different from driving a browser with CDP?
For Chromium apps it does not replace CDP; launch --cdp opens a verified DevTools port so frameworks like Playwright or Puppeteer drive the web contents, while native menus, dialogs and windows stay on the accessibility path. The accessibility route is what covers native apps that CDP cannot reach.
Does agent-desktop work on Mac?
Yes. The README's installation instructions and permission model are macOS-specific, and it requires macOS 13.0 or newer. Screenshots additionally require Screen Recording permission, and the Notification Center opener requires Automation permission for System Events.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/lahfir-agent-desktop)