browser-control: A Rust CLI for Agent-Driven Browser Automation
A tiny, fast Rust CLI that drives a real browser over the Chrome DevTools Protocol — built for coding agents.
At a glance
- What is it?
- browser-control is a single static Rust binary that exposes a real browser as a set of shell commands over the Chrome DevTools Protocol. It is designed specifically for AI coding agents, not for human-driven test scripts.
- Who is it for?
- browser-control is the right tool for an AI coding agent that needs to drive a real browser from a shell pipeline with no SDK wrapper and no long-running server to manage. It is not designed for human-authored test suites where fixtures, assertions, and reporters are expected out of the box.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 35 days ago.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What problem browser-control solves for AI agents
Most browser automation frameworks were designed for human developers writing test scripts. They bundle APIs for waiting, assertion, and fixture management, and they assume a programmer reads the output. browser-control starts from a different assumption: the consumer is an AI coding agent running shell commands in a loop.
The project's stated design is a browser as a debuggable pipe. Every capability is a subcommand that prints compact text or JSON to stdout. The agent reads that output, decides what to do next, and issues another command. There is no SDK, no language-specific client library, and no long-running server process to manage beyond the browser itself. The README states the design goal directly: if an agent can run a shell command, it can drive a browser.
At version 0.1.1, released on 2026-06-21, browser-control is early-stage software. The crate on crates.io is named browser-control-cli (the name browser-control was already taken), but the installed binary is called `browser-control`.
Stable element refs: snapshot replaces brittle selectors
The central design choice that distinguishes browser-control from lower-level CDP clients is its snapshot command. Running `browser-control snapshot` returns a numbered list of actionable elements on the current page, each assigned a stable short handle: @e1, @e2, @e3, and so on. A click command then takes `@e2` as its target rather than a CSS path.
The README describes this as refs built for LLMs. A language model cannot reliably write correct CSS selectors for pages it has never seen, but it can read a list of element handles and act on them by number. The handles remain stable across repeated snapshot calls on the same page state, so an agent that receives @e3 can click it immediately without recalculating a selector.
CSS selectors and pixel coordinates are still supported when they are more appropriate. The three target styles available to every action command are: a snapshot ref (the default), a CSS selector with an optional --wait polling timeout, and x,y pixel coordinates for canvas elements or mapped regions. The eval command wraps a top-level return automatically, so expressions like `eval 'const x = 1; return x'` work without a function wrapper.
Installing and starting the first session
The quickest install path uses cargo:
cargo install browser-control-cli
browser-control --versionThis builds from crates.io and places the `browser-control` binary on your PATH. Prebuilt binaries for specific platforms are available from the GitHub Releases page. To download and install the macOS ARM64 binary without a Rust toolchain:
curl -fsSL https://github.com/keon/browser-control/releases/latest/download/browser-control-aarch64-apple-darwin.tar.gz | tar xz
install -m 0755 browser-control /usr/local/bin/To build from source with a reproducible toolchain:
git clone [email protected]:keon/browser-control.git
cd browser-control
rustup toolchain install
cargo build --locked --releaseThe release binary lands at `target/release/browser-control`. The build requires Chrome or Chromium reachable over CDP. The `launch` subcommand starts a browser instance for you. A typical first session:
browser-control init
browser-control launch https://example.com
export BROWSER_CONTROL_CDP_URL=http://127.0.0.1:9222
browser-control doctor
browser-control snapshot`init` creates the .browser-control/ workspace directory. `doctor` reports the detected endpoint, browser process, daemon status, and workspace. Running snapshot after that prints the numbered element list for the current page. If you already have a browser listening on a CDP port, skip `launch` and set `BROWSER_CONTROL_CDP_WS` or `BROWSER_CONTROL_CDP_URL` directly.
The self-observing daemon and CDP escape hatch
browser-control runs a hidden daemon that maintains in-memory rings for DOM events, network requests, and console output. The `browser-control events` and `browser-control network` commands read from these rings. When a command fails, the tool writes a compact trace to .browser-control/traces/ so an agent can diagnose the failure without human intervention.
For anything that the helper surface cannot express directly, browser-control exposes two low-level commands. `eval` executes arbitrary JavaScript in the page context. `cdp` sends raw Chrome DevTools Protocol messages. The README is explicit that these prevent hitting a wall and needing to switch tools: an agent always has a path to the underlying protocol.
The press command handles keyboard input, including combinations such as `ctrl+a`, `cmd+shift+t`, and clipboard operations. The README notes that it maps editing combos to renderer editing commands explicitly to handle the macOS case where app menu interception would otherwise swallow keystrokes in headless mode.
For cloud deployments, browser-control supports remote browsers over CDP. The same commands work against local Chrome, or against managed browser providers such as Browser Use, Steel, Hyperbrowser, or Browserbase by setting the appropriate CDP URL in the environment variable.
Reproducible builds and the pinned toolchain
browser-control's build is intentionally locked down. The Cargo.toml lists every direct dependency as an exact version rather than a range. The Cargo.lock freezes transitive dependencies. The rust-toolchain.toml pins the Rust compiler version, and `rustup toolchain install` reads it automatically.
The README recommends always building with --locked in CI:
cargo build --locked
cargo test --locked
scripts/verify.shThis means two engineers building from the same commit get the same binary, and a CI server's output is comparable to a local build. The static binary has no runtime dynamic linking dependency on the system, which simplifies deployment in container environments. The minimum Rust edition is 2021 and the minimum rust-version declared in Cargo.toml is 1.92.
Where browser-control is the wrong fit, and how it compares to Playwright
browser-control v0.1.1 is explicitly early-stage. The README carries no claims of feature completeness. There is no assertion library, no test reporter, no fixture management, and no built-in mechanism for parallelizing sessions. Teams building integration test suites that require those features and that are maintained by human engineers should look at Playwright instead.
Playwright is Microsoft's browser automation framework, written primarily for TypeScript and Python test authors. It bundles its own browser binaries, includes a full assertion library, generates visual traces, and has a large ecosystem of plugins and integrations. Playwright's programming model assumes a developer writes the test script. browser-control's model assumes a language model reads the snapshot and writes the command.
The distinction matters: Playwright's API is expressive for a programmer but verbose for an LLM that needs to issue one compact command at a time. browser-control's subcommand interface produces output a language model can parse directly. The choice between them depends on whether a human or an AI agent is the primary author of the browser interaction.
Editorial conclusion
browser-control is the right tool for an AI coding agent that needs to drive a real browser from a shell pipeline with no SDK wrapper and no long-running server to manage. It is not designed for human-authored test suites where fixtures, assertions, and reporters are expected out of the box. Before adopting it, verify that your target platform can install the browser-control binary and that the Chrome or Chromium version in your environment supports the CDP endpoints the tool requires; run `browser-control doctor` to confirm the connection before building anything on top of it.
Frequently asked questions
How does browser-control differ from running raw CDP commands directly?
browser-control adds a snapshot layer that generates stable @e1/@e2 element refs so an agent can click elements by handle rather than by CSS selector. It also maintains an in-memory event daemon and writes failure traces to .browser-control/traces/ automatically. Raw CDP calls are still available via the `cdp` subcommand when needed.
Does browser-control require a specific browser to be installed?
Yes. The tool requires Chrome or Chromium reachable over the Chrome DevTools Protocol. The `browser-control launch` command will start a local browser instance, and `browser-control doctor` reports whether the endpoint, browser process, and daemon are detected correctly.
Can browser-control connect to a cloud browser provider instead of a local Chrome?
Yes. Set BROWSER_CONTROL_CDP_WS or BROWSER_CONTROL_CDP_URL to the remote CDP endpoint. The README lists Browser Use, Steel, Hyperbrowser, and Browserbase as supported providers, all using the same subcommand interface as a local session.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/keon-browser-control)