CLI tool
trycua/cua avatar
trycua/cua

trycua/cua: background computer-use drivers, sandboxes and benchmarks

Scale computer-use 2.0 with open-source drivers, cross-OS fleets, and benchmarks for training, evaluation, and data generation.

26,358 stars1,828 forksHTMLMIT

At a glance

What is it?
The trycua/cua monorepo ships four things under one MIT licence: a background driver that does not steal the cursor, an ephemeral sandbox SDK, a benchmark runner, and Lume for macOS VMs. The hard part is knowing which one you actually need.
Who is it for?
Adopt the driver if you need an agent to operate native desktop apps on a machine a human is also using, and adopt the sandbox SDK if you want the same API across Linux, macOS, Windows and Android. Skip it if you need a stable public API surface: the repository publishes nightly Lume builds and frequent computer-server releases, and the README itself points at an open issue where the Sequoia unattended preset still stops at the Accessibility step.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 4 days ago.
What is it written in?
Mainly HTML, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

Four products in one repository, and the choice you have to make first

The README opens with a path selector rather than a description, and that is the honest framing. trycua/cua is not one tool. It is a driver for background desktop control, a sandbox SDK called Cua, a benchmark and RL environment runner called Cua-Bench, and Lume, a macOS virtualisation layer built on Apple's Virtualization.Framework. The repository layout agrees: libs/ holds the packages, cua-bench sits at the top level as a runnable project, and clusters/, infra/ and nix/ hold the deployment and build scaffolding.

The audience is narrower than the tagline suggests. This is for people building or evaluating computer-use agents, meaning software that observes a screen and emits clicks, keystrokes and gestures. If you want a chat assistant that answers questions, nothing here helps you. If you want an agent that fills in a desktop application, or a harness that scores such an agent against OSWorld and ScreenSpot, the four components cover that whole loop.

The decision you have to make before writing code is whether your agent runs on a machine a person is also using. Cua Drivers exists for that case. The README states it drives native desktop apps in the background and that agents click, type and verify without stealing the cursor or focus. The sandbox path assumes the opposite: a disposable VM or container that nobody else is touching.

How the sandbox SDK abstracts four operating systems behind one Python API

The Cua package is the part with the clearest mechanism in the README. A single context manager selects the operating system, and the object handed back exposes the same names regardless of what is underneath. The README example calls Sandbox.ephemeral(Image.linux()) and then uses sb.shell.run, sb.screenshot, sb.mouse.click, sb.keyboard.type and sb.mobile.gesture. Swap Image.linux() for .macos(), .windows() or .android() and the call sites do not change.

That uniformity is the design claim, and the support table is where it gets tested. Linux containers, Linux VMs, macOS, Windows and Android are all marked as available both in the cloud at cua.ai and locally through QEMU. Bring-your-own-image, listed as .qcow2 and .iso, is marked available locally and as coming soon in the cloud. So the local path is currently the more complete one, which is the reverse of what a hosted-first product usually looks like.

The README does not document what happens when an OS lacks a capability the shared interface exposes. The example calls sb.mobile.gesture on a Linux sandbox in the same block, which reads as illustrative rather than literal. Treat the surface as uniform and the behaviour as OS-specific until you check the API reference at cua.ai/docs/cua/reference/agent-sdk.

Installing cua and running a first sandbox command

The Python SDK installs from PyPI. The README shows a single pip command, and the inline comment in the example pins the floor at Python 3.11 or later. The workspace pyproject.toml is stricter, declaring requires-python of >=3.12 and <3.14, so a 3.11 interpreter may work for the published package while the monorepo itself will not resolve.

bash
pip install cua

The first real use is an ephemeral Linux sandbox that runs a shell command and takes a screenshot. Save this as a file and run it; the README presents it as the canonical starting point.

python
# Requires Python 3.11 or later
from cua import Sandbox, Image

async with Sandbox.ephemeral(Image.linux()) as sb:
    result = await sb.shell.run("echo hello")
    screenshot = await sb.screenshot()
    await sb.mouse.click(100, 200)
    await sb.keyboard.type("Hello from Cua!")

What you should see is the shell result object from echo hello and a screenshot of the sandbox desktop. The click and type calls act on that desktop, so if the screenshot comes back black or the click lands nowhere, the problem is the image, not your agent code. The README points at cua.ai/docs/cua/guide/get-started/set-up-sandbox for the full setup and cua.ai/docs/cua/examples for more.

Cua Drivers and the background-input problem on Wayland

The driver is the component with the sharpest technical constraint. Its purpose is to let an agent operate native applications without taking the cursor or focus, which is what makes it usable on a workstation someone is actively using. The README says the same CLI and MCP server work across macOS, Windows and Linux, and names Claude Code, Cursor, Codex and OpenClaw as clients.

Linux is where the README gets specific instead of promotional. It states that Linux supports X11 and compositor-specific Wayland routes, with explicit limits for raw background input. That phrasing matters: on Wayland, background input depends on which compositor you run, and the project documents the limits rather than claiming parity. If your fleet is Wayland-based, the compositor is the first thing to check, not the last.

Installation is a shell one-liner on macOS and Linux and a PowerShell one-liner on Windows, both fetched from cua.ai rather than from a package registry. The README then says to follow the post-install instructions, which is doing real work in that sentence: the interesting configuration lives after the script, not inside it. Architecture notes and an optional agent skill pack are in libs/cua-driver/README.md, which the top-level README treats as the source of truth for this component.

What Lume does, and the one preset that is not verified end to end

Lume creates and manages macOS and Linux VMs on Apple Silicon through Virtualization.Framework, and the README claims near-native performance. It is the piece that makes the macOS row of the sandbox support table reachable on hardware you own.

The README walks through a vanilla macOS install. You fetch an Apple restore image with lume ipsw, create a VM with lume create and the --unattended flag, then run it with lume run. The --unattended option prepares the guest offline, and the built-in sequoia and tahoe presets create the lume user, enable SSH, configure autologin, and disable sleep and screen locking. Default credentials are lume / lume, which is convenient for a throwaway VM and obviously wrong for anything reachable from a network.

The README is unusually direct about reliability here. It says the Tahoe flow is E2E verified, and that Sequoia may still open the Accessibility step of Setup Assistant on its first display boot, linking issue #2155. That is a real limitation stated in the project's own words, and it is the kind of thing that turns an unattended provisioning script into an interactive one. If your pipeline depends on Sequoia being hands-off, that issue is the first thing to read.

Cua-Bench, and why the benchmark runner is the least self-contained part

Cua-Bench evaluates computer-use agents on OSWorld, ScreenSpot, Windows Arena and custom tasks, and exports trajectories for training. The README's quickstart clones the repository, changes into cua-bench, installs with uv tool install -e ., creates a base image with cb image create linux-docker, and then runs a dataset with cb run dataset datasets/cua-bench-basic --agent cua-agent --max-parallel 4.

Note the shape of that sequence. It is not a pip install. You are cloning the monorepo, installing the package in editable mode with uv, and materialising a Docker image before the first run. That is a heavier commitment than the sandbox SDK, and it is the right weight for what it does: running a benchmark means controlling the environment precisely, which is why the base image is an explicit step rather than something the tool fetches for you.

The --max-parallel 4 flag is worth understanding before you raise it. Parallelism here means multiple agent-driven sandboxes running at once, so the ceiling is set by the machine's CPU, memory and disk rather than by the tool. The README does not discuss resource sizing for that flag, and the CLI reference at cua.ai/docs/cuabench/reference/cli-reference is where a reader would have to look.

Alternatives, and the trade-off against a single-OS tool

The obvious alternative for anyone whose problem is desktop automation rather than agent evaluation is a mature single-platform automation library such as Playwright for browsers or a native accessibility-API binding for one desktop OS. The difference in approach is not quality, it is scope. Those tools assume a specific surface and expose it in full detail. Cua assumes four surfaces and exposes a shared subset, which is why the README can show one code block covering Linux, macOS, Windows and Android.

If you only ever target Linux containers, that abstraction is a cost with no benefit. You are reading documentation about macOS restore images and Wayland compositors to run a container. A container-native automation stack will be simpler and better documented for that case.

The reverse holds too. If your evaluation needs to compare the same agent across operating systems, the shared API is the whole point, and stitching together four platform-specific libraries means writing the adapter layer yourself. The README's support table is the honest summary of where that bet currently pays off: local QEMU covers every column including bring-your-own-image, while the cloud covers everything except BYOI.

Licence, release cadence and what an upgrade actually costs

The repository is MIT licensed, and the workspace pyproject.toml repeats that as license = { text = "MIT" }. MIT is permissive: it allows commercial use and modification with attribution and no warranty. That is a statement about the licence text, not advice about your situation; the Lume path bundles Apple restore images and runs Apple's Virtualization.Framework, so the terms that govern the guest OS are separate from the terms that govern this code.

The release pattern is the practical cost. The listed releases include computer-server v0.3.45 and v0.3.43 on the same day, plus a nightly Lume build tagged nightly-lume-v0.5.4-nightly.20260828. Nightly tags mean the Lume surface moves faster than a version number suggests, and the last push to the repository was on 2026-08-28.

Version bumps are not done by hand. The Makefile states that version bumps are managed via GitHub Actions workflows and points at release-bump-version.yml, and its own targets are dry runs: make dry-run-patch-core, make dry-run-minor-core, make dry-run-major-core, plus make show-versions to print current versions from the per-package .bumpversion.cfg files. For a consumer of the packages, the consequence is that the workspace uses uv with a members list under tool.uv.workspace covering agent, core, computer, computer-server, som, mcp-server and bench-ui. Pinning versions is your job, not the project's.

Editorial conclusion

Adopt the driver if you need an agent to operate native desktop apps on a machine a human is also using, and adopt the sandbox SDK if you want the same API across Linux, macOS, Windows and Android. Skip it if you need a stable public API surface: the repository publishes nightly Lume builds and frequent computer-server releases, and the README itself points at an open issue where the Sequoia unattended preset still stops at the Accessibility step. Verify three things before committing: that your Wayland compositor is one of the routes cua-driver documents, that your host meets the Python 3.12 to 3.14 range in the workspace pyproject.toml, and that the driver install script from cua.ai/driver/install.sh is acceptable to your security review, since it is fetched over the network rather than installed from a package index.

Frequently asked questions

What is a cua agent?

In this repository, cua-agent is the AI agent framework for computer-use tasks, listed in the packages table and documented at cua.ai/docs/cua/reference/agent-sdk. It is the component that decides what to click or type; the sandbox SDK and driver are what execute those actions.

What's a cua driver?

Cua Drivers is the component that drives native desktop apps in the background on macOS, Windows and Linux, so agents click and type without stealing the cursor or focus. It exposes the same CLI and MCP server across all three, and source documentation lives in libs/cua-driver/README.md.

What are CuA projects?

The term covers the family of tools in this repository: Cua Drivers for background desktop control, Cua for agent-ready sandboxes, Cua-Bench for benchmarks and RL environments, and Lume for macOS virtualisation. The README presents them as four paths chosen by what you are building.

What is cuadriver?

The README writes the name as Cua Drivers, and the package directory is libs/cua-driver. It is installed with a shell script on macOS and Linux or a PowerShell script on Windows, both hosted at cua.ai, followed by post-install instructions.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/trycua-cua.svg)](https://hysenlabs.com/projects/trycua-cua)
Community notes

Community notes