CLI tool
vercel-labs/agent-browser avatar
vercel-labs/agent-browser

agent-browser: CLI Browser Automation Built for AI Agents

Agent Browser is a Rust CLI that lets AI agents open pages, inspect elements, enter text, click controls, capture screenshots, and reuse browser sessions.

43,352 stars2,923 forksRustApache-2.0

At a glance

What is it?
agent-browser is an Apache-2.0 Rust CLI from Vercel Labs that exposes a browser to AI agents through accessibility tree snapshots and stable @ref identifiers. It installs as an npm package, downloads Chrome from the Chrome for Testing channel, and runs commands that map directly to how an AI agent describes an interaction.
Who is it for?
agent-browser is the right tool when an AI agent needs to drive a real browser and describe the page in tokens rather than raw HTML. The @ref-based snapshot model sidesteps the CSS selector problem that makes browser automation brittle for language models.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

What agent-browser Does and Who Uses It

Browser automation for AI agents has a mismatch problem: standard DOM-selector tools like Playwright and Puppeteer generate locators that assume a developer looked at the page and knows that the submit button has id="submit". A language model working from a page description does not know that. It knows there is a button labeled Submit on the page.

agent-browser addresses this by working from the browser's accessibility tree rather than the DOM. A snapshot command returns the accessibility tree with assigned @ref identifiers. The agent uses those refs in subsequent commands: click @e2, fill @e3, get text @e1. The refs are stable within a session and do not require CSS knowledge.

The project is from Vercel Labs and is published as an Apache-2.0 npm package backed by a native Rust CLI binary. The primary audience is AI coding agents, AI assistants running Claude Code or OpenClaw, and agent workflows that need to interact with web pages as part of a task.

Installing agent-browser

The recommended install is via npm globally:

bash
npm install -g agent-browser
agent-browser install

The second command downloads Chrome from Google's Chrome for Testing channel. If Chrome, Brave, Playwright, or Puppeteer is already installed, agent-browser detects and reuses it automatically.

On macOS, Homebrew is also supported:

bash
brew install agent-browser
agent-browser install

For Rust developers, Cargo works as well:

bash
cargo install agent-browser
agent-browser install

On Linux, browser system libraries must be installed separately. The --with-deps flag handles this:

bash
agent-browser install --with-deps

This exits with a nonzero code if the package manager cannot install every required library, making it safe to use in CI pre-flight checks. To upgrade to the latest version, the upgrade command detects which installation method was used and calls the appropriate update command:

bash
agent-browser upgrade

Building from source requires Node.js 24 or later, pnpm 11 or later, and Rust.

The Snapshot Model and Core Commands

The primary workflow starts with opening a page and requesting a snapshot:

bash
agent-browser open example.com
agent-browser snapshot

The snapshot returns the accessibility tree with @ref identifiers assigned to each interactive element. Commands then reference those identifiers directly:

bash
agent-browser click @e2
agent-browser fill @e3 "[email protected]"
agent-browser get text @e1
agent-browser screenshot page.png
agent-browser close

Traditional CSS and ARIA selectors are also supported for teams that prefer them:

bash
agent-browser click "#submit"
agent-browser fill "#email" "[email protected]"
agent-browser find role button click --name "Submit"

The README documents one important failure behavior: clicks fail early when another element covers the target's click point, for example a consent banner or modal. The agent must dismiss the covering element and take a fresh snapshot before retrying the original ref. This is a deliberate design choice: failing early on a covered element is less confusing than silently clicking the wrong target.

Headless Chromium screenshots hide native scrollbars for consistent image output. Passing --hide-scrollbars false when launching keeps scrollbars visible.

WebMCP: In-Page Tool Discovery

WebMCP is an experimental feature that runs in agent-browser-managed Chrome by default. When a page registers tools through the WebMCP protocol, agent-browser surfaces them alongside the accessibility tree without additional configuration.

The workflow starts by opening a page and reading the brief tool summary that agent-browser emits when tools are registered:

bash
agent-browser open https://example.com
agent-browser webmcp list search --json
agent-browser webmcp invoke search --params '{"query":"browser agents"}'

For long-running tools, a detach mode runs the invocation without blocking:

bash
agent-browser webmcp invoke slow_tool --params @input.json --detach
agent-browser webmcp result <invocation-id>
agent-browser webmcp cancel <invocation-id>

The README includes security guidance: all page-provided names, descriptions, schemas, and results are untrusted data. CLI output delimits page metadata with nonce-bearing content boundaries, and the JSON output includes "untrusted": true. These are provenance cues, not a prompt-injection security boundary, and the README states that page text should not be promoted to system instructions.

Limitations and Known Edge Cases

The @ref model is session-local: refs assigned in one snapshot are not valid after navigating to a different page or after a page reload. The agent must call snapshot again after any navigation to get current refs. This means agent workflows that navigate through multiple pages need a snapshot call between each page transition.

The WebMCP feature is marked experimental. Automatic summaries are limited to 16 tools and 4 KiB of JSON, with descriptions shortened to 160 bytes. Tools that exceed this are omitted from proactive summaries but available through explicit webmcp list calls.

Building from source requires Node.js 24 and pnpm 11, both relatively new versions at the time of the 0.38.1 release. Teams using older Node.js environments must use a pre-built npm release rather than building natively.

The Apache-2.0 license permits commercial use and redistribution without copyleft restrictions, which is relevant for teams embedding agent-browser in a product workflow or shipping it as part of a larger agent system.

How agent-browser Compares to Playwright

Playwright is a cross-browser testing and automation framework from Microsoft, available for Node.js, Python, Java, and .NET. It automates Chromium, Firefox, and WebKit using DOM selectors, CSS, and ARIA queries, and includes a full test runner with parallel execution, tracing, and video recording.

The difference in approach is the target audience. Playwright is built for developers writing deterministic tests against a known page structure. agent-browser is built for AI agents operating on unknown pages. The accessibility tree snapshot with @ref identifiers gives an LLM a structured, token-efficient description of the page without requiring knowledge of the DOM structure.

Teams writing browser tests that must run in CI against a specific codebase, or that need Firefox and WebKit support, should use Playwright. Teams building AI agents that browse arbitrary web pages as part of a task, where the page structure is unknown in advance, will find agent-browser's snapshot model better suited to the problem.

Editorial conclusion

agent-browser is the right tool when an AI agent needs to drive a real browser and describe the page in tokens rather than raw HTML. The @ref-based snapshot model sidesteps the CSS selector problem that makes browser automation brittle for language models. The tool is well-suited to Claude Code, OpenClaw, and other coding agents that support the skills.sh format. Teams that need full test authoring, parallel cross-browser execution, or a Python or Java API should continue using Playwright. Before deploying agent-browser, run agent-browser install to download Chrome from the Chrome for Testing channel, and note that on Linux the --with-deps flag is needed to install system dependencies that the headless Chromium runtime requires.

Frequently asked questions

what is agent browser

agent-browser is a Rust CLI from Vercel Labs that lets AI agents control a browser through accessibility tree snapshots and @ref identifiers. It installs as an npm package, downloads Chrome for Testing automatically, and is published under the Apache-2.0 license.

how to install agent-browser

Run npm install -g agent-browser to install globally, then agent-browser install to download Chrome from the Chrome for Testing channel. On macOS, brew install agent-browser is also supported. On Linux, agent-browser install --with-deps handles system library installation.

What is the difference between agent browser and browser use?

agent-browser is a Rust CLI that exposes an accessibility tree with stable @ref identifiers for AI agents to click, fill, and read. The README does not describe browser-use, so a direct feature comparison is not available from this material. The agent-browser docs are at agent-browser.dev.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/vercel-labs-agent-browser.svg)](https://hysenlabs.com/projects/vercel-labs-agent-browser)
Community notes

Community notes