All concepts
Concept

What is Computer use?

Computer use is the practice of letting an AI model drive a real browser or desktop session, clicking, typing and reading the screen on your behalf. A browser agent applies the same idea to web pages only, usually through a browser automation layer.

Published September 28, 2026

How computer use actually works

A computer-use agent is a loop. The model receives a goal in natural language, observes the current state of a screen or page, chooses one action, and repeats until the goal is met or it gives up. The observation step is what separates implementations. Some agents read the accessibility tree, a structured list of roles, names and states that a browser exposes for assistive technology. Others take a screenshot and reason over pixels. Some do both, using the tree for reliable element references and the screenshot as a fallback.

The action step is narrower than it looks. In practice an agent needs a small vocabulary: open a URL, click a coordinate or an element reference, type text, press a key, scroll, wait, capture a screenshot. Everything else is composition. vercel-labs/agent-browser is described as a Rust CLI that lets agents open pages, inspect elements, enter text, click controls, capture screenshots and reuse browser sessions, which is close to that minimum vocabulary exposed as commands. feder-cr/AIHawk runs a real browser from a prompt, either standalone on 127.0.0.1:8765 or as an MCP server for Claude Code, Codex and Gemini CLI, so the same loop can be driven by a chat client instead of a bespoke script.

The connection layer matters as much as the loop. The Model Context Protocol (MCP) has become a common way to expose browser actions as tools that a coding assistant can call. AIHawk ships an MCP server. seleniumbase/SeleniumBase includes an MCP server for agent-driven browsing alongside its Python test framework. citrolabs/ego-lite takes a different angle: it is an MIT-licensed browser aimed at agent automation, with a dedicated Space per agent and the ego-browser skill as the connection layer, so an agent works inside a browser profile that already holds your logged-in state.

Finally there is the sandbox question. Running an agent against your everyday desktop risks stray clicks and leaked credentials. trycua/cua ships, under one MIT licence, a background driver that does not steal the cursor, an ephemeral sandbox SDK, a benchmark runner, and Lume for macOS VMs. The background driver and the sandbox SDK address different problems: the first keeps a human usable while the agent works, the second gives the agent a disposable machine. The README does not document rollback for actions taken inside a sandbox, so treat cleanup as something you configure rather than something the tool guarantees.

When you need it, and when you do not

You need computer use when the target has no usable API and the workflow is genuinely visual or stateful. Logging into a vendor portal with a session cookie, filling a multi-step form that depends on earlier answers, or extracting a figure that only appears after a chart renders are all cases where an HTTP client returns nothing useful. Stagehand, from browserbase/stagehand, is built for this end of the problem: it is a TypeScript, Python and Go SDK that wraps Playwright-style APIs with natural language actions like act, observe and extract, and it targets production browser agents rather than test scripts.

You do not need it when a documented API exists. A REST or GraphQL call is faster, cheaper, easier to test and far easier to debug than a model deciding where to click. You also do not need it for plain crawling. apify/crawlee is an Apache-2.0 TypeScript library for building crawlers in Node.js, with a persistent request queue, pluggable storage and proxy rotation, and it works with Puppeteer, Playwright, Cheerio, JSDOM and raw HTTP. If the job is fetching and parsing pages at volume, Crawlee is the cheaper layer; a model in the loop adds latency and nondeterminism for no benefit.

A middle case is test automation. SeleniumHQ/selenium is the reference implementation of the W3C WebDriver spec, covering Java, JavaScript, Python, Ruby, .NET and the Grid. If you already have a Selenium suite, adding an agent on top of the same stack is often less work than adopting a separate runtime. SeleniumBase wraps Selenium 4.49.0 and pytest with a stealth layer called CDP Mode, plus an MCP server. That combination suits Python teams already writing tests, and it is a poor fit for anyone who wants a minimal browser client.

The honest rule: use computer use when the interface is the only interface, and use it on a copy of the account or machine where a wrong click is recoverable.

Pitfalls and limits you will hit

Nondeterminism is the first limit. A model choosing actions can take a different path on two identical runs, which makes failures hard to reproduce. Stagehand's self-healing primitives trade some determinism for adaptability, which is a deliberate choice: the agent recovers when a selector changes, but the exact sequence is no longer fixed.

Element references are the second limit. Accessibility-tree refs are stable within a page load and fragile across deployments. A ref that pointed to a submit button may point elsewhere after a redesign, and a coordinate click is worse, because it depends on viewport size, zoom and scroll position. vercel-labs/agent-browser is described as fast, but its design choices around Chrome for Testing and accessibility-tree refs shape who should adopt it: teams that can pin a browser build and tolerate ref churn get the speed, teams that cannot should look at selector-based wrappers.

Platform coverage is the third limit, and it is usually stated plainly. citrolabs/ego-lite is MIT-licensed and runs on macOS today; its README is explicit that Windows and Linux are still on the roadmap. feder-cr/AIHawk ships Windows and Linux builds only, and installing it means uvx plus an OpenRouter key. Neither project claims full cross-platform parity, so check your target OS before committing.

Maturity and maintenance are the fourth. Maintenance status has to be read from each repository's archived flag and last push date, not from enthusiasm. NanmiCoder/cc-haha is MIT licensed and its last push was on 2026-08-22; it wraps Claude Code in a cross-platform desktop app with multi-session workspaces, worktree launches, diff review, GUI permission approval, provider switching and Computer Use. apify/crawlee is pushed frequently, but the README does not document rollback and the v4 line is still at a release candidate, so pin a version you have exercised.

Finally, permissions and credentials. An agent that can click can also approve a dialog, delete a record or accept a payment. Put it on a scoped account, keep destructive actions behind a human approval step, and prefer sandboxes over your primary session.

How it shows up in open-source projects

The projects below fall into four groups, and the grouping is more useful than a ranked list.

Browser clients and agent runtimes. vercel-labs/agent-browser is a Rust CLI for opening pages, inspecting elements, entering text, clicking controls, capturing screenshots and reusing sessions. feder-cr/AIHawk runs a real browser from a prompt, standalone on 127.0.0.1:8765 or as an MCP server for Claude Code, Codex and Gemini CLI. citrolabs/ego-lite is an MIT-licensed browser with a dedicated Space per agent and the ego-browser skill as the connection layer, macOS only for now.

SDKs and frameworks. browserbase/stagehand is a TypeScript, Python and Go SDK that wraps Playwright-style APIs with natural language actions like act, observe and extract, aimed at production browser agents. SeleniumHQ/selenium is the W3C WebDriver reference implementation across Java, JavaScript, Python, Ruby, .NET and the Grid. seleniumbase/SeleniumBase wraps Selenium 4.49.0 and pytest with CDP Mode, a stealthy Chromium layer, plus an MCP server.

Infrastructure and data generation. trycua/cua is a monorepo shipping four things under one MIT licence: a background driver that does not steal the cursor, an ephemeral sandbox SDK, a benchmark runner, and Lume for macOS VMs. apify/crawlee is an Apache-2.0 TypeScript crawler library with a persistent request queue, pluggable storage and proxy rotation.

Desktop and workspace tools. Hammerspoon/hammerspoon is an MIT-licensed macOS app that exposes system APIs to a Lua scripting engine, so window management, hotkeys and app control live in one init.lua file; it is not click-to-configure, and that is the point. NanmiCoder/cc-haha is a cross-platform desktop workspace for Claude Code and AI coding workflows, with multi-agent sessions, branch worktree support, diff review, permissions, model switching, Computer Use, a skill marketplace and messaging integrations; its last push was on 2026-08-22.

Read each project's own README for the details this page does not cover, because several of them do not document rollback, cost controls or cross-platform parity.

How to choose between them

Start from the interface you are allowed to touch, not from the model. If an API exists, use it. If you must drive a page, decide whether you need a full agent that plans, or a scripted automation that only occasionally calls a model for a fragile step.

If you are writing Python tests, SeleniumBase and Selenium are the natural stack, and SeleniumBase adds CDP Mode and an MCP server. If you are in Node.js and the job is crawling at volume, Crawlee is the cheaper layer and its proxy rotation and persistent queue are the features that matter. If you want an agent to operate inside your logged-in browser on macOS, ego-lite is explicit about that use case and about its current platform limit. If you need sandboxes or macOS VMs for training and evaluation, trycua/cua is the only project in this list that names those pieces, and the hard part is knowing which of the four you actually need. If you want a chat client to drive a browser through MCP, AIHawk and SeleniumBase both expose one. If you want a desktop surface with permissions and diff review around an agent, cc-haha is the closest match, with its last push on 2026-08-22.

None of these projects removes the need for a scoped account, a pinned browser build and a human approval step on destructive actions.

In practice

Computer use is a loop that observes a screen and takes one action at a time, and its value depends on whether the interface you need has an API. Read the README of whichever project matches your stack, check its archived flag and last push date, and try it on a scoped account before pointing it at anything you cannot undo.

vercel-labs/agent-browserAgent Browser is a Rust CLI that lets AI agents open pages, inspect elements, enter text, click controls, capture screenshots, and reuse browser sessions.43,352 stars · RustSeleniumHQ/seleniumA browser automation framework and ecosystem.34,509 stars · Javafeder-cr/AIHawkOpen-source AI browser agent for web automation: a web browsing agent and computer-use agent in plain English. Browser MCP for Claude Code and Gemini CLI.31,627 stars · Pythontrycua/cuaScale computer-use 2.0 with open-source drivers, cross-OS fleets, and benchmarks for training, evaluation, and data generation.26,358 stars · HTMLapify/crawleeCrawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.25,821 stars · TypeScriptbrowserbase/stagehandThe SDK For Browser Agents. Most existing browser automation tools either require you to write low-level code in a framework like Selenium, Playwright, or Puppeteer, or use high-level agents that can be unpredictable in production.25,387 stars · TypeScriptcitrolabs/ego-liteThe fastest browser for AI agents to run browser automation, built for sharing your logged-in browser state with your AI agents, like Codex or Claude Code, without disturbing you. Zero cost, zero config.16,679 stars · JavaScriptHammerspoon/hammerspoonStaggeringly powerful macOS desktop automation with Lua16,195 stars · Objective-CNanmiCoder/cc-hahacc-haha is a cross-platform desktop workspace for Claude Code and AI coding workflows, with multi-agent sessions, branch worktree support, diff review, permissions, model switching, Computer Use, skill marketplace, and remote team-friendly messaging integrations.14,714 stars · TypeScriptseleniumbase/SeleniumBaseBrowser automation, web scraping, and testing. CDP Mode provides a Stealthy Chromium instance that bypasses bot-detection systems. Includes an MCP server.13,046 stars · Pythonoctalmage/robotjsNode.js Desktop Automation. 12,777 stars · Cgo-vgo/robotgoRobotGo, Go Native cross-platform RPA, GUI automation, Auto test and Computer use @vcaesar10,845 stars · Go

Sources

  1. vercel-labs/agent-browser repository
  2. SeleniumHQ/selenium repository
  3. feder-cr/AIHawk repository
  4. trycua/cua repository
  5. apify/crawlee repository