CLI tool
sleepinginsummer/agent-browser-cli avatar
sleepinginsummer/agent-browser-cli

agent-browser-cli: a Chrome extension plus a Rust daemon, with three disagreeing versions

使用 agent-browser-cli 进行浏览器感知与控制。适用于标签页扫描/切换、页面 JS 执行、Cookie、CDP、contentSettings、截图、文件上传、下拉框点击、tmwd_cdp_bridge 初始化和 Web 工具排障

626 stars72 forksRustMIT

At a glance

What is it?
agent-browser-cli gives an agent control of the Chrome session you already have, login state and cookies included, by driving a Manifest V3 extension rather than launching its own browser. The detail that trips people up is bookkeeping: the README, the npm manifest and the Cargo manifest each declare a different version, and the npm scope drops a letter from the GitHub owner.
Who is it for?
Use agent-browser-cli if your agent work depends on pages that need a real login state, since reusing your own Chrome profile is the whole reason it exists rather than launching a clean browser. Do not pick it as a Playwright replacement, because the README says plainly that it is not one.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 62 days ago.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Three files, three version numbers, and a misspelled scope

The project information section of the README states the current version as 0.3.5.

The npm manifest says 0.3.7, which matches the newest release tag.

The Cargo manifest says 0.3.4.

So the Rust crate is three patch releases behind the published npm package, and the README's own version line is two behind. Nothing in the repository synchronises the Cargo version with the npm version: there is a script that syncs the optional platform dependency versions, and it runs as an npm lifecycle hook on a version bump that also stages the manifest and the lockfile, but it only touches the JavaScript side.

There is a second mismatch in the naming. The GitHub account is `sleepinginsummer`, and the npm scope is `@sleepinsummer`. The install command, the uninstall command and all five platform packages use the shorter spelling.

None of this stops the tool working, but it means the version in the README is not a reliable way to check what you installed.

It is not Playwright, and the login state is the point

The README draws the distinction itself: this project is not Selenium or Playwright, and it is more suited to helping an agent read pages and act precisely inside an existing browser session.

That is a design statement about what gets reused. Rather than launching its own browser, this one connects through a Manifest V3 extension loaded into your real Chrome, so the session keeps its login state and its cookies.

The capability list has nine entries. Scan the current Chrome tabs for page title, URL and tab id. Switch to a specified tab, reusing the existing page and login state. Open a new tab to a target URL. Execute JavaScript in the page to read the DOM, forms, state and page data. Read the current page's cookies. Call Chrome's DevTools Protocol for lower-level control. Take screenshots. Upload a local file to a page file input. Operate dropdowns, buttons and forms.

The README also credits its lineage: the browser control capability was extracted and adapted from GenericAgent, including its web driver, its page-simplification work and the bridge extension itself. The listed improvements over that origin are avoiding a fresh browser connection on every command, a startup lock so concurrent CLIs cannot double-bind the port, and a reworked attach and detach flow that reduces page flicker and stray attachments after an abnormal exit.

Elements are addressed as handles, not as selectors

The command list is where the agent-facing design becomes clear. Instead of CSS selectors, elements are referred to by short handles: click takes `@e1`, fill takes `@e2` followed by text, key presses go through a send-keys command with a target handle, and a mouse click also takes a handle.

So the loop is: inspect, get handles, act on handles. That is what makes an agent's actions survive a page re-render, because a handle identifies an element the tool has already resolved rather than a string the agent has to reconstruct.

The rest of the surface is grouped by concern. Discovery commands are `tabs`, `scan` and `snapshot`, and a raw `exec` takes arbitrary JavaScript such as a document title expression. Capture commands are `screenshot` with a full-page flag, and `save-pdf`. Network and console have start and list pairs. Tab opening takes a window-and-focus flag or a group-title flag, and there is a short alias command at the end of the list.

One more decision is worth naming. The tree listing truncates URLs and omits a session key by default specifically to reduce token use, with a `--full` flag when the complete fields are needed. That is a token budget written into the default output rather than left to the caller.

Two ports, and three ways to change one of them

Two fixed numbers carry the whole architecture. Port 18765 is the extension's WebSocket port, which the Chrome extension connects to and which can be changed with a dedicated set-extension-port command. Port 18767 is the CLI's HTTP API port, used for session reuse, and it explicitly cannot be used as the extension port.

Changing the extension port has three documented paths, and they do not behave the same way. The first is the command itself:

bash
agent-browser-cli set-extension-port 18766

That writes the configuration file and, if the daemon is running, restarts the daemon automatically so the new port takes effect immediately.

The second path is editing the configuration by hand. The file lives at `~/.agent-browser-cli/config.json`, is generated if it does not exist, and holds a single key for the extension port. But a manual edit is not picked up until you run a restart command, because the daemon is already listening on the old number.

The third is the extension's own popup, which can change the port and reconnect immediately.

In all three cases the extension port must match the value in the CLI configuration, so the popup is a convenience rather than a second source of truth.

Profile labels are aliases, and the CLI refuses ambiguous ones

With several Chrome profiles or browser instances open, the internal profile and browser identifiers are long, so the tool supports short labels that you set per profile and then pass with a profile flag.

The commands are straightforward: look up a tab or a browser by identifier, set a label for a profile, then list tabs filtered by that label.

Two behaviours are specified precisely. A label is only an alias; internal routing still uses the browser, profile and tab identifiers as a three-part key. And if a label matches more than one profile inside the current daemon, the CLI reports the ambiguity rather than picking one.

The README then recommends setting labels through the CLI rather than through the extension popup, and the reason is a validation difference rather than a convenience one. The CLI checks that a label is unique across profiles within the running daemon. The popup is described as a local convenience entry that does not guarantee that uniqueness.

So there is a documented way to create an ambiguous label, and the tool tells you which entry point to avoid.

Five prebuilt platforms, provenance on, and no Windows on ARM

The npm package is a launcher. The `bin` entry points at a Node script, `engines` requires Node 18 or newer, and the real work is done by a Rust binary delivered as an optional dependency chosen per platform.

There are five of those packages, all at the same version: darwin on arm64, darwin on x64, linux on x64, linux on arm64, and win32 on x64. Windows on arm64 is the gap.

Two details in the manifest are worth noting. Publishing is configured with provenance enabled, which is the npm supply-chain attestation path. And the version lifecycle script regenerates the optional dependency versions and then stages the manifest and lockfile, so those five version strings are kept in step with the package version automatically.

The Rust dependency list is small and legible. The HTTP client has default features off and TLS through rustls rather than a system OpenSSL, so building does not need a C toolchain. One crate handles Windows path canonicalisation, which is the long-standing problem of Windows paths that begin with an unusual prefix. Another is a file-locking library, which is what sits behind the startup lock that stops two CLIs binding the same port. The web server has JSON and WebSocket support, the HTML parser pair is what powers the page-simplification scripts, and the rest is a CLI parser, a UUID generator and the async runtime.

WSL needs mirrored networking, and page dialogs are deliberately left alone

The platform requirements are more specific than a support matrix usually is.

Windows, macOS and Linux are all supported, and Windows includes WSL. The WSL path needs WSL 2.0.0 or newer, and it recommends enabling mirrored networking mode under Windows 11 22H2 or newer so that WSL can reach a Chrome bridge service running on the host's localhost. Without that, the two halves of the architecture are on different network stacks and cannot see each other. Linux additionally requires the local Chrome or Chromium build to support installing extensions.

There is also a runtime requirement that is easy to miss: Chrome needs at least one ordinary web page tab open. A tab parked on the browser's own settings page or on a blank page is not enough, because the extension attaches to page content.

The dialog behaviour is a deliberate restraint. The extension no longer overrides the page's own alert, confirm and prompt functions by default. Native dialogs are suppressed only while a CLI page-script command is running, and restored afterwards, specifically to avoid permanently interfering with a business page's global functions.

The on-page badge follows the same restraint: it is draggable, expands on hover, hides itself after ten seconds without a command, and returns after roughly three hundred seconds when the service disconnects and reconnects.

Editorial conclusion

Use agent-browser-cli if your agent work depends on pages that need a real login state, since reusing your own Chrome profile is the whole reason it exists rather than launching a clean browser. Do not pick it as a Playwright replacement, because the README says plainly that it is not one. Before installing, check that your platform has a prebuilt package, since five exist and none covers Windows on ARM, and read the port configuration rather than assuming the default applies, because the extension port and the CLI port are different numbers that cannot be swapped.

Frequently asked questions

install agent browser cli

Run `npm install -g @sleepinsummer/agent-browser-cli` and then `agent-browser-cli tabs`, or build from source with `cargo build --release` and run the binary in `target/release`. You also have to load the Chrome side: download `chrome-extensions.zip` from the latest release, extract it, enable developer mode in `chrome://extensions`, load the unpacked `tmwd_cdp_bridge` directory, and make sure at least one real web page tab is open.

What is the agent-browser used for?

The README lists eight uses: automation testing against a real browser session with form submission and login-state pages, frontend page debugging, style debugging, collecting data from pages that need a logged-in session, scripting browser operations, letting an agent drive admin backends and low-code platforms, page structure analysis, and security research or reverse-engineering assistance. It describes page style work as an aid for debugging rather than a complete design tool.

Does agent-browser-cli replace Playwright?

No. The README states that this project is not Selenium or Playwright and is aimed at helping an agent read pages and act precisely inside an existing browser session. The difference is that it attaches to your real Chrome through an extension, so the login state and cookies of that profile are already there.

Why does agent-browser-cli need a Chrome extension?

Because the CLI does not launch a browser. It connects through the Manifest V3 extension in `assets/tmwd_cdp_bridge`, which is why installation includes loading an unpacked extension and keeping a real web page tab open. Two ports are involved: 18765 for the extension's WebSocket and 18767 for the CLI's HTTP API, and the extension port must match `extension_port` in `~/.agent-browser-cli/config.json`.

Official sources

  1. Issues
  2. License: MIT
  3. README
  4. Releases
  5. sleepinginsummer/agent-browser-cli on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/sleepinginsummer-agent-browser-cli.svg)](https://hysenlabs.com/projects/sleepinginsummer-agent-browser-cli)