# Browser4: a CDP-native browser engine built to be driven by agents

> A Kotlin engine that pairs a Rust CLI with snapshot-based interaction, SQL over stored HTML, and a verifier-style split between doing the work and checking it, aimed at large scraping runs.

**platonai/Browser4** — Browser4 — an AI-native browser engine for autonomous agents, intelligent extraction, and large-scale web automation.

- Repository: https://github.com/platonai/Browser4
- Website: https://browser4.io
- Stars: 1,152 · Forks: 152
- Language: Kotlin
- License: Apache-2.0
- Published: 2026-10-07 · Updated: 2026-10-07 · Language: en
- Canonical page: https://hysenlabs.com/projects/platonai-browser4

## A Kotlin engine with a Rust CLI and a Chrome extension on the side

The repository layout tells you how the pieces divide. `browser4-core/`, `browser4-boot/`, `browser4-rest/` and `browser4-agentic/` are the Kotlin service side, packaged with Maven into a jar. `cli/` is the Rust command line binary that drives it. `cdp-protocol/` is a first-party Chrome DevTools Protocol implementation, which is the detail that explains the architecture: rather than wrapping an existing CDP client, this project speaks the protocol itself.

That is why the README can claim a coroutine-safe engine designed for very high visit counts, and why the module count is as large as it is. Two more directories widen the surface: `chrome-extension/` and `coworker/`, alongside `browser4-plugins/`, `browser4-coding/`, `skills/` and `docs-dev/`. Apache-2.0 licensed, with the README available in English and Chinese and a Gitee mirror.

The Kotlin and TypeScript mix in the topic list, plus `browser4-coding/`, points at what the maintainers call a programming-agent kernel with more than fifty `coding.*` tools for agents that are building Browser4 artifacts, or Browser4 itself. That is a self-development story rather than a feature you consume, and it is worth setting expectations early on what this is: an engine with an agent-facing surface, not a library you import.

## Snapshot with refs is the interaction model everything else builds on

Browser4 does not ask you to write selectors first. The documented flow is `snapshot -i --boxes`, which returns clickable and typeable elements with short refs such as `e15`, and then you act on those refs with `click`, `fill`, `select`, `hover`, `drag`, `scroll` and `wait`. Every interaction command also accepts CSS selectors, so you can start with refs and harden into selectors later.

The README's own flow is worth reading as a template, since it shows the whole shape in seven lines:

```bash
browser4-cli open --headed https://example.com/login
browser4-cli snapshot -i --boxes
browser4-cli fill e3 "user@example.com"
browser4-cli fill e4 "secret" --submit
browser4-cli wait --load networkidle
browser4-cli snapshot -i
```

Note the `--headed` flag. The README notes humans usually want to see the browser, and `batch` exists to chain multiple steps so you are not paying process startup per action.

The iframe handling is the part that separates this from a thin Playwright wrapper. `frames` lists the frame tree, `frame "<iframe selector>"` scopes subsequent element commands into that frame, and `frame main` returns to the main document. Same-origin iframes are supported fully, and the README claims this removes the need for manual `contentDocument` evaluation, which is exactly the escaping problem that trips up payment forms and embedded editors.

## Extraction comes in two kinds, and only one of them costs tokens

The tool selection guide in the README is a decision tree, and the branch that matters is whether the extraction needs a language model. `htmlsnapshot get text` and `htmlsnapshot get all text` pull from stored HTML snapshots with selectors and spend nothing. `htmlsnapshot query --sql @query.sql` runs X-SQL against a snapshot for correlated fields across repeated items. `eval --json` handles live JavaScript and complex DOM logic. `extract` takes natural language such as find the product price, and that one needs an LLM key.

Above that sits WebMiner, which the README describes as ML clustering that turns HTML corpora into spreadsheet and report views with no LLM tokens. There is also a progressive experience store that reuses learned selectors and blockers across runs, so a workflow that worked yesterday does not start from zero today.

For volume there are `crawl` and `swarm`. A list of known URLs in a file goes to `crawl --seed-file urls.txt --depth 0 --sql @query.sql`; a start URL you want followed goes to `crawl <url> --out-l`. The `--sql` flag is what makes output structured instead of scraped.

If you are choosing an approach, the practical rule is that deterministic extraction should stay on X-SQL and CSS selectors, and the model should only be reached for when the page resists both. That keeps cost proportional to the hard pages rather than to the volume.

## Running it yourself: Docker with Mongo, and an honest note on privileges

The README leans on a skill file rather than a conventional install guide, telling you to point your agent at `https://browser4.io/SKILL.md`. There is also a DeepSeek Harness plugin path with two commands:

```bash
dsh plugin --profile web add dsh-browser4                  # npm registry
dsh plugin --profile web add github:platonai/dsh-browser4  # GitHub
```

The self-hosted route is the Docker Compose file at the repository root, which runs the `galaxyeye88/browser4:latest` image with an embedded `mongo:latest` on port 27017 and publishes the service on 8182. The environment block is where the tuning lives: `BROWSER_CONTEXT_MODE` has a DEFAULT setting, `BROWSER_CONTEXT_NUMBER` and `BROWSER_MAX_OPEN_TABS` only apply when the mode is not DEFAULT, `BROWSER_DISPLAY_MODE` is set to HEADLESS, and `SERVER_ADDRESS` is `0.0.0.0`.

One Dockerfile comment is worth more than most of the settings. Headless Chrome renders through shared memory, and the Docker default of 64MB is too small for multiple tabs, which makes renderer processes crash or wedge and shows up as `Runtime.evaluate` hanging. The compose file sets `shm_size: 1g` for exactly this reason, and also raises `nofile` to 65535.

The Dockerfile builds a standalone jar with Maven and copies the finished `Browser4.jar` into place, deliberately avoiding a `COPY` that would pull a stale jar from the host build context. That is the kind of detail that suggests the container path has been debugged rather than sketched.

## Where this differs from Playwright, Puppeteer and Browserless

The honest comparison is against browser automation libraries, and the difference is in what is bundled. Playwright and Puppeteer give you a driver and a browser. You write the extraction, you decide what a selector means, and if the page changes shape you own the breakage. Browser4 instead ships the decision tree, the stored HTML snapshots, the SQL layer and the learned-selector store, on the theory that at scale the extraction logic is the expensive part to maintain.

Browserless is the closer comparison because it also runs a browser as a service. The difference is that a hosted browser endpoint gives you a connection to a page, while this project gives you storage over the page as well, which is what makes `htmlsnapshot` and X-SQL possible at all. If your workload is a few thousand pages, that storage is overhead. If it is hundreds of thousands with the same shapes repeated, it is the reason to switch.

There is a real cost to the bundling. This is a Kotlin service with an embedded MongoDB, a Maven build, a Rust CLI and a Chrome extension, not a `npm install`. Your operational surface is a running server rather than a library linked into your test process, and the README's own framing, including a hosted cloud at browser4.io and docs at a separate site, suggests a product rather than a utility.

One limitation to name: the project publishes to Maven Central as `ai.platon.pulsar/browser4-core`, while the npm package is `browser4-cli`, so the two halves land in different registries under different group names. Budget for that when you set up internal mirrors.

## Release cadence, license, and what to read first

The release list is busy and current: v4.14.0-rc.6 on 2026-09-17, v4.13.19 on 2026-09-16, v4.13.18 on 2026-09-12, with the last push on 2026-09-20. A release candidate landing the day after a patch release is normal for this kind of project, but it means pinning matters more than usual.

The README is where to start and also where it stops. It carries the capability list, the tool selection guide, the interaction flow, an architecture section and a modules overview, then defers the rest to the documentation site. It also carries a proxy configuration section for unblocking website access, which tells you the maintainers consider network egress a solved-enough problem to document rather than a blocker.

Two things the README does not settle. It claims 100k to 200k complex page visits per machine per day through swarm and crawl scale-out, but publishes no benchmark data, no hardware profile and no methodology, so treat that as a target from the project rather than a measurement. And `examples/` carries a single `browser4-examples/` directory with a `pom.xml`, which is a thin example surface for a project this size.

The `test-fixture-server` called MockSite is the more useful thing in the repository for your own work, since it gives you a controlled target to validate extraction against instead of pointing the crawler at a live site.

## Conclusion

Browser4 is worth a look when your page work is repetitive enough to need refs, batches and stored HTML rather than a script that re-parses live DOM, and when you have a machine to run it on. The performance figure in the README, 100k to 200k complex page visits per machine per day, is the project's own claim about its swarm and crawl scale-out and is not independently verified in the repository. Two things to check before committing: the 4.x release train is still shipping release candidates, with v4.14.0-rc.6 on 2026-09-17 against v4.13.19 the day before, so pin a version; and the Docker path wants `--privileged` with an embedded MongoDB on port 27017, which is a heavier footprint than a headless Chrome script. Start with `browser4-cli open --headed https://example.com/login` and a `snapshot -i --boxes` call before building anything around it.

## FAQ

### How do I install Browser4 and run a first browser automation task?

The README points agents at the skill file at browser4.io/SKILL.md, which handles installing or upgrading `browser4-cli`. For self hosting, the repository ships a `docker-compose.yml` that runs the `galaxyeye88/browser4:latest` image with an embedded MongoDB and publishes the service on port 8182. A first interactive run is `browser4-cli open --headed https://example.com/login` followed by `browser4-cli snapshot -i --boxes`.

### What is Browser4 used for compared with Playwright or Puppeteer?

Playwright and Puppeteer are drivers where you write the extraction yourself. Browser4 bundles the extraction layer instead: stored HTML snapshots, X-SQL queries, an ML clustering step and a store that reuses learned selectors across runs. That is more to operate, since it is a Kotlin service with an embedded MongoDB, but it moves the maintenance burden from your scripts to the engine when you are processing pages at volume.

### Does Browser4 need an LLM API key to run?

Only for the paths that use one. The `extract` command takes natural language and needs an LLM key, while `htmlsnapshot get text`, `htmlsnapshot get all text` and `htmlsnapshot query --sql` are deterministic and are described as costing no tokens. WebMiner clustering is also described as producing spreadsheet and report views with no LLM tokens. You can set an OpenAI-compatible endpoint through the Docker environment, for example `DEEPSEEK_API_KEY`.

### How does Browser4 handle pages with iframes?

Through built-in frame switching rather than manual evaluation. The `frames` command lists the frame tree, `frame "<iframe selector>"` scopes later element commands into that frame, and `frame main` returns to the main document. Same-origin iframes are fully supported, and the README says this avoids the `contentDocument` evaluation that the usual approach needs.

## Sources

- [License: Apache-2.0](https://github.com/platonai/Browser4/blob/main/LICENSE)
- [platonai/Browser4 on GitHub](https://github.com/platonai/Browser4)
- [Project website](https://browser4.io)
- [README](https://github.com/platonai/Browser4/blob/main/README.md)
- [Releases](https://github.com/platonai/Browser4/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/platonai-browser4
