Model or dataset
browser-act/skills avatar
browser-act/skills

BrowserAct Skills: Browser Automation CLI Built for AI Agents

Browser automation CLI built for AI agents. Break through anti-bot walls, hand off to humans across platforms when stuck. Parallel multi-task execution, independent multi-session operation, isolated multi-account browsing.

6,012 stars305 forksPythonMIT

At a glance

What is it?
BrowserAct Skills is a Python CLI tool that gives AI agents a real browser with stealth fingerprinting, CAPTCHA solving, human handoff, and parallel session isolation, designed from the ground up for agent-driven automation rather than human-scripted workflows.
Who is it for?
BrowserAct Skills is worth evaluating if you are running AI agents that need to reach anti-bot-protected pages, manage multiple authenticated sessions in parallel, or hand off to a human when the agent is stuck. The free tier covers most automation needs; the paid tier adds managed proxies and stealth browsers beyond the first five.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 37 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What BrowserAct Skills Solves for AI Agent Workflows

Standard browser automation libraries such as Playwright or Selenium are built for human-authored test scripts. AI agents present different requirements: they need compact output that fits in a token budget, indexed interaction that requires no DOM parsing, semantic memory to match browsers to tasks by meaning, and a way to recover when the agent hits a challenge it cannot solve.

BrowserAct Skills addresses each of these. The README lists four core requirements a browser for agents must meet: breaking through anti-bot blocks across three progressive layers (environment fingerprinting, execution-layer CAPTCHA solving, and human escalation), supporting three distinct browser modes for different real-world scenarios, providing zero-interference concurrency across parallel sessions, and producing compact indexed output rather than raw HTML or JSON.

The tool operates as a command-line interface that agents call with shell commands. It works with any agent that can execute shell commands and load Skills, including Claude Code, Cursor, VS Code, OpenCode, OpenClaw, Codex and Gemini CLI.

Three Browser Modes and How They Differ

The README documents three browser modes matched to distinct use cases. The `chrome` mode reuses the user's local Chrome login state, either by importing a profile or by attaching via CDP. This is the right mode for workflows that depend on an existing authenticated session without re-entering credentials.

The `stealth` privacy mode runs a fresh fingerprint and empty profile per session, with proxy rotation and zero residue after the session ends. This targets batch scraping tasks that do not require login state and where each run should leave no trace.

The `stealth` fixed identity mode uses a stable fingerprint and stable IP for maintaining logged-in accounts across a parallel set of browsers. This is designed for multi-account workflows where each account needs its own isolated session identity and consistent behavior to avoid triggering bot detection.

The three modes map to real scenarios: a shopping agent reusing your login, a data collection pipeline running anonymously at scale, and a social media management tool keeping five accounts isolated and stable.

Installing BrowserAct Skills and First Commands

The README instructs agents to install by telling an AI agent directly:

> Install browser-act. Skill source: https://github.com/browser-act/skills/tree/main/browser-act . Verify it works after installation.

For direct command-line use after installation:

bash
browser-act stealth-extract https://example.com

This extracts content from a protected page with zero configuration. For full browser automation:

bash
browser-act --session my-task browser open <id> https://example.com
browser-act --session my-task state
browser-act --session my-task click 3
browser-act --session my-task input 2 "hi"

The `state` command returns an indexed list of interactive elements on the page. The `click 3` and `input 2 "hi"` commands interact by index, which eliminates DOM parsing. At the start of each session, agents run:

bash
browser-act get-skills core --skill-version 2.0.2

This returns the current environment state, browser list and command reference in a single call. The requirements.txt lists `requests>=2.28.0` and `python-dotenv>=1.0.0` as the only Python dependencies.

Anti-Bot Layers, CAPTCHA Solving and Human Handoff

BrowserAct Skills describes a three-layer approach to anti-bot defenses. The environment layer applies stealth fingerprint spoofing and TLS rotation to avoid triggering most blocks before any interaction occurs. For sites that do trigger additional challenges, the execution layer provides `solve-captcha` for automatic CAPTCHA resolution and `stealth-extract` for pulling protected page content in a single command.

For cases where automation cannot proceed, the human layer provides `remote-assist`. This generates a live URL that the user can open on any device to take over the browser session manually. When the user finishes the manual step, the agent continues from where it left off. The README describes this as "cross-platform remote handoff," meaning the user can open the link on a phone, tablet or different computer.

This escalation path addresses a real gap in agent automation: agents fail on multi-step authentication flows, visual puzzles, or pages that require human judgment. Rather than letting the agent fail silently or loop, `remote-assist` turns the failure into a hand-off.

Concurrency, Session Isolation and Security Gating

The README describes two concurrency models. Cross-browser parallel execution runs independent sessions with separate cookies, fingerprints and proxies, so sites cannot correlate them. Same-browser multi-session execution shares a login state across sessions but keeps task execution independent, meaning tasks do not block each other.

Session ownership and explicit naming prevent conflicts when multiple agents run simultaneously. Each session is identified by name; a second agent cannot accidentally interact with a session owned by a different agent.

The security model includes a confirmation-gating requirement for sensitive operations. Browser creation and deletion, profile import, proxy changes, and security and privacy toggles all require explicit user approval. Prior approvals do not carry over to subsequent operations. The README states this is "enforced at the Skill layer, not a configuration toggle," which means it cannot be disabled through configuration.

What Is Free and What Costs Money

The README documents a three-tier access model. Everything is free without signup for Chrome and Chrome-direct browser automation. Features that require free account login include stealth browsers up to a limit of five, `stealth-extract`, `solve-captcha`, `remote-assist`, privacy mode and Skill Forge. The paid tier adds stealth browsers beyond five and managed dynamic or static proxies.

The practical implication is that most agent automation tasks fall within the free tier. The cost boundary appears when running more than five concurrent stealth sessions or when needing managed residential or datacenter proxies for reliability at scale.

BrowserAct also ships Skill Forge, described as a personal scraping engineer feature that builds and tests a reusable scraping bot from a description, then runs it from the cloud. This is part of the cloud-managed execution option and requires a login.

Limitations and Cases Where BrowserAct Skills Is Not the Right Tool

BrowserAct Skills is designed for agent-driven automation, not for human-authored test suites or CI pipeline browser testing. The indexed interaction model and compact text output are optimized for LLM reasoning, not for assertion-based test frameworks. Teams that need browser testing integrated with a testing framework should use Playwright or Selenium directly.

The tool runs on Windows, macOS and Linux but stealth headless mode is a specific variant noted in the README; standard headless may behave differently on some anti-bot systems. The documentation for specific anti-bot success rates is not given numerically in the README.

The last push to the repository was on 2026-08-24. The project does not have formal GitHub releases. The MIT licence permits commercial use without restrictions.

Editorial conclusion

BrowserAct Skills is worth evaluating if you are running AI agents that need to reach anti-bot-protected pages, manage multiple authenticated sessions in parallel, or hand off to a human when the agent is stuck. The free tier covers most automation needs; the paid tier adds managed proxies and stealth browsers beyond the first five. Teams with simple, non-blocked scraping needs may find standard browser automation tools sufficient. Before adopting it, verify that the `browser-act stealth-extract` command reaches the specific sites you need, and review the confirmation-gating requirement, which enforces user approval for sensitive operations such as browser creation, profile import and proxy changes.

Frequently asked questions

How do you install BrowserAct Skills in Claude Code?

According to the README, tell your AI agent: "Install browser-act. Skill source: https://github.com/browser-act/skills/tree/main/browser-act . Verify it works after installation." The agent handles the installation steps. Full installation details are in the docs/installation.md file in the repository.

How do you use BrowserAct Skills in Claude Code?

After installation, agents call BrowserAct commands from the shell. Run `browser-act get-skills core --skill-version 2.0.2` at the start of a session to load environment state and the command reference. Then use `browser-act --session <name> state` to inspect a page, and `browser-act --session <name> click <index>` to interact.

What AI agents does BrowserAct Skills work with?

The README lists Claude Code, Cursor, VS Code, OpenCode, OpenClaw, Codex and Gemini CLI as compatible agents. It states that BrowserAct works with any agent that can execute shell commands and load Skills.

Official sources

  1. browser-act/skills on GitHub
  2. Issues
  3. License: MIT
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/browser-act-skills.svg)](https://hysenlabs.com/projects/browser-act-skills)