# Tencent/BrowserSkill: let an AI agent borrow your logged-in browser

> BrowserSkill is a CLI plus Chromium extension that hands shell-capable agents a visible, isolated Agent Window built on the login state you already have. It is a solved install and a partly unsolved trust problem.

**Tencent/BrowserSkill** — Let AI agents use your real, logged-in browser without interrupting your work. CLI + extension for browser automation across any shell-capable AI agent.

- Repository: https://github.com/Tencent/BrowserSkill
- Stars: 7,961 · Forks: 562
- Language: TypeScript
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/tencent-browserskill

## The problem: agents that cannot reach the pages that matter

Most browser automation assumes a clean profile. You launch a fresh Chromium, it has no cookies, and every page behind a login is either scripted with stored credentials or skipped. That works for public pages and breaks for the internal ones. The pages an engineer actually wants an agent to touch are behind SSO, behind a corporate VPN, behind a session that took a hardware key to establish.

BrowserSkill attacks that specific gap. The README states the goal plainly: agents work with sites you are already signed into, without separate test accounts. The target user is not a QA team running a grid. It is one developer with Cursor or Claude Code open in one window and a browser full of logged-in tabs in another, who wants the agent to read or click something without taking the browser away. The README's framing is that the agent must borrow a tab explicitly, return it when the task is done, and leave the rest of the browser alone. That is a product decision as much as a technical one: the workflow being protected is the human's own browsing.

## How the bsk daemon and the extension split the work

There are two local runtime pieces, and the repository layout confirms the split. The Rust workspace has two crates, bsk-cli and bsk-protocol, with the protocol crate generating a schema through a `dump-schema` binary that `cli:build` runs before the TypeScript build. The extension is a separate pnpm package, @browser-skill/extension, with its own build, dev, test and zip scripts. So the CLI is Rust, the extension is TypeScript, and the schema in bsk-protocol is the contract between them.

The Cargo manifest shows the transport: tokio and tokio-tungstenite are workspace dependencies, which points to a WebSocket link between the daemon and the extension. The daemon also carries reqwest with rustls, flate2, tar and zip, which is what an updater needs when it downloads and unpacks a release. The README describes the agent-facing side as a CLI: any agent that can call a shell can use BrowserSkill through `bsk`, with no lock-in to a model or framework. The extension side is what actually holds browser permissions, and it is distributed through the Chrome Web Store and Edge Add-ons rather than loaded from a local build in the normal path.

The visible Agent Window is the interesting part of the architecture. Rather than driving your foreground tab, tasks run in a separate window, which is what makes the claim about not interrupting your work possible. The README does not document how tab identity is tracked across that boundary, so treat the borrow-and-return contract as a documented promise rather than a mechanism you can inspect from the README alone.

## Installing bsk and running a first agent task

The README offers two paths. The recommended one is to paste a single instruction into a shell-capable agent, which then follows AGENT_INSTALL.md to install the CLI and skill and walks you through loading the extension. The manual path is three steps: CLI, extension, skill.

On macOS or Linux the installer script places the binary in `~/.local/bin`:

```bash
curl -fsSL https://raw.githubusercontent.com/Tencent/BrowserSkill/main/install.sh | sh
```

On Windows the PowerShell installer targets the same directory:

```powershell
irm https://raw.githubusercontent.com/Tencent/BrowserSkill/main/install.ps1 | iex
```

Confirm the binary is on your PATH before going further. The README gives this check:

```bash
bsk --version
```

Next, install the extension. Chrome users go through the Chrome Web Store listing, Edge users through Edge Add-ons; on other Chromium-based browsers the README says to use the Chrome Web Store build. Firefox is listed as planned, not supported.

Then teach your harness how to call `bsk`:

```bash
bsk install-skill
```

The command opens a selector. Use Space to pick the harness and Enter to install. `bsk install-skill --list` prints internal variants and install paths. If you maintain your own instructions, point at the file explicitly:

```bash
bsk install-skill --harness cursor --source ./SKILL.md
```

An explicit `--source` stays custom even when its contents match the bundled skill, and existing installations are skipped unless you add `--force`. After that, the README's own test is a tab you already have open: ask the agent to borrow it, do something with it, and give it back.

## The skill update rules are stricter than they look

Most tools overwrite their own generated files on upgrade and leave you to notice. BrowserSkill does the opposite, and the README spends real space on it. Daemon startup, `session start` and `doctor` update managed skills only when the installed contents still match the last installed version. If you edited SKILL.md, automatic updates pause and your edit survives.

The edge cases are where this gets fiddly. An older installation with no content baseline is enrolled automatically only if it exactly matches the current bundled skill, and enrollment writes the source marker without rewriting SKILL.md. A file with an unrecognized source marker, a local edit, or a differing historical file produces a WARN from `doctor` with the reason and recovery options. That warning does not fail the health check: `--json` reports `status: "warn"` and `ok: true`. A concurrent install or sync is reported as deferred and retried later.

This is a sensible design for anyone who treats the skill file as their own. It is also a design that can leave you on an old skill without a loud failure, because a WARN with `ok: true` is easy to scroll past. If you want to keep your current instructions as a deliberate customization, the README's route is `bsk install-skill --harness cursor --source <existing-SKILL.md> --force`.

## Sandboxed agents need a different setup

The README flags one environment where the default install misbehaves: agent sandboxes that reap background processes after each command. The daemon is a long-lived process, so a sandbox that kills everything between commands will kill it too. The documented fix is to keep the daemon in a persistent host environment and connect from the sandbox using a shared `BSK_HOME` plus `BSK_AUTO_START=0`, with the details in docs/sandboxed-agents.md. Ordinary local use keeps automatic startup on.

That is a real constraint, not a footnote. It means BrowserSkill is not a drop-in for a containerized agent runner that treats each command as a fresh process. You have to decide where the daemon lives and make that location reachable, which is an infrastructure decision the README pushes to a separate document. If your agent runs in a sandbox you do not control, check that document before assuming the one-line install is enough.

## Where BrowserSkill is the wrong tool

The honest limitation is the flip side of the feature. BrowserSkill exists to drive a browser that holds your real credentials, in a real profile, on your real machine. That is exactly the property you do not want in a shared CI runner. If your task is a nightly check against a staging site, a headless Chromium in a container is simpler, has no extension to install, no daemon to keep alive, and no risk of an agent clicking something in a session that belongs to a person.

The platform matrix is the second boundary. macOS on Apple Silicon and Intel, Linux on x64 and ARM64, and Windows x64 are listed. Chrome and Edge are supported; other Chromium-based browsers are expected to work when they support unpacked Chromium extensions. Firefox is planned, which is a polite way of saying it does not work yet. If your team standardizes on Firefox, this is not the tool.

The human-in-the-loop feature is a limitation in disguise. The README says that when a task hits a captcha, login, confirmation dialog, or another human-only step, the agent can ask you to take over and then continue afterwards. That is good design for a present operator and useless for an unattended run. A pipeline that assumes no one is watching will stall at the first captcha.

## How it differs from Playwright and Puppeteer

Playwright and Puppeteer are libraries. You write a script, the script launches or connects to a browser, and the script owns the session. They are excellent at what they do, and they are the right answer when the browser is disposable.

BrowserSkill inverts the ownership. The browser is yours, already running, already authenticated, and the agent is a guest that must request a tab and hand it back. The interface is a shell command rather than an API call, which is what makes it harness-agnostic: the README lists Cursor, Claude Code, Codex, OpenClaw, CodeBuddy, WorkBuddy, Pi, Hermes Agent and DeepSeek Harness, and the claim is that anything with shell access qualifies. You are not writing a script per site; you are giving an existing agent a capability it did not have.

The cost of that inversion is determinism. A Playwright script either passes or fails in a known way. An agent borrowing your tab is doing something you did not script, in a profile you care about, and the README's answer to that is the borrow-and-return contract plus a visible Agent Window rather than a sandbox. If you need reproducible runs, stay with the libraries.

## Licence, releases and what upgrades cost

The repository is MIT licensed, and the Cargo workspace declares `license = "MIT"` at the workspace level. MIT is permissive: you can use, modify and redistribute the code, including in closed products, provided the copyright notice and permission notice are preserved. That is a summary of the licence text, not legal advice; read LICENSE in the repository before relying on it for anything commercial.

Upgrade cost is low but not zero. Releases are versioned separately, and the recent tags make the split visible: ext-v0.2.1 for the extension and cli-v0.2.1 for the CLI, both dated 2026-09-09, with ext-v0.2.0 a week earlier. Because the extension ships through browser stores, extension updates arrive on the store's schedule rather than yours, while the CLI updates through its own installer. If the two drift apart, bsk-protocol is the schema that has to stay compatible, which is presumably why it exists as its own crate. The repository's last push was on 2026-09-10, one day after those releases.

The skill file adds a second upgrade surface. Because automatic updates pause when you edit SKILL.md, a customized skill can sit unchanged across CLI upgrades. That is the intended behaviour, and it means `bsk doctor` is the thing to run when you are unsure which version of the instructions your agent is actually reading.

## Conclusion

Adopt BrowserSkill if your agent work depends on sessions you cannot recreate (internal dashboards, SSO-gated tools) and you are willing to run a local daemon plus a store-installed extension on Chrome or Edge. Do not adopt it if you need Firefox, if you need a headless CI browser with no human present, or if you cannot accept an agent driving a tab that holds real credentials. Verify three things before committing: that `bsk --version` reports the version you expect, that `bsk doctor` returns a clean status rather than a WARN, and that the tab-borrow and tab-return behaviour matches what the README promises on your own sites.

## FAQ

### What is the best browser agent?

That depends on whether the browser is disposable. If it is, a library such as Playwright or Puppeteer gives you a script you control. BrowserSkill targets the opposite case: an agent that needs the browser you are already logged into, driven through the `bsk` CLI in a separate Agent Window.

### What are browser automation skills?

In BrowserSkill the skill is a set of instructions that teaches an agent harness how to call the `bsk` CLI. It is installed with `bsk install-skill`, which opens a harness selector, and it is updated automatically only while its contents still match the last installed version.

### How do I install the agent browser skill for BrowserSkill?

Run `bsk install-skill`, use Space to select the harness, and press Enter. To install your own instructions instead, use `bsk install-skill --harness cursor --source ./SKILL.md`; an explicit source stays custom even if its contents match the bundled skill, and existing installations are skipped unless you add `--force`.

### How do I use the agent browser skill with BrowserSkill?

Once the CLI, the extension and the skill are in place, ask your agent to borrow a tab you already have open, complete the task, and return it. The README's stated contract is that the agent borrows the tab explicitly, returns it when done, and leaves the rest of your browser alone.

## Sources

- [Issues](https://github.com/Tencent/BrowserSkill/issues)
- [License: MIT](https://github.com/Tencent/BrowserSkill/blob/main/LICENSE)
- [README](https://github.com/Tencent/BrowserSkill/blob/main/README.md)
- [Releases](https://github.com/Tencent/BrowserSkill/releases)
- [Tencent/BrowserSkill on GitHub](https://github.com/Tencent/BrowserSkill)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/tencent-browserskill
