# ntegrals/openbrowser: a TypeScript agent framework that drives Playwright from a task string

> Open Browser wraps Playwright in an agent loop with OpenAI, Anthropic and Google models, ships a CLI and a REPL, and adds a sandbox package for resource-limited runs. The trade-off is a young project with no releases and a README that stops short of several operational details.

**ntegrals/openbrowser** — Let AI agents browse the web. An autonomous toolkit for browser-based AI agents.

- Repository: https://github.com/ntegrals/openbrowser
- Stars: 9,536 · Forks: 865
- Language: TypeScript
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/ntegrals-openbrowser

## The problem Open Browser targets

Writing a Playwright script for a known page is straightforward. Writing one that survives a page you have not seen is not, because selectors, layout and flow all change. Open Browser takes the position that the agent should decide the steps at runtime. You hand it a sentence such as the README's example, "Find the top story on Hacker News and summarize it", and the framework handles clicking, typing, scrolling and extraction.

The intended reader is a TypeScript developer who wants browser automation that tolerates variation, and who is willing to pay per-token for that tolerance. The repository is a monorepo with three packages: open-browser for the core agent logic, browser control, DOM analysis and model integration; @open-browser/cli for the command line; and @open-browser/sandbox for resource-limited execution. If your task is a fixed login-and-download job that runs nightly, this is more machinery than the job needs.

## How the agent loop is put together

The README describes an Agent constructed from a viewport, a model and a task, with a settings object. The viewport comes from createViewport, the model from createModel, and the agent runs until it finishes or hits a limit. That shape tells you where the control points are: the browser is a separate object you can configure or reuse, and the model is a separate object you can swap.

Settings govern the loop rather than the browser. stepLimit caps iterations, commandsPerStep caps actions inside one iteration, and failureThreshold stops the run after a number of consecutive failures. enableScreenshots controls whether page images go into the agent context, and contextWindowSize sets the token budget for the conversation. allowedUrls and blockedUrls restrict navigation. Cost tracking and stall detection are listed as features, though the README does not explain how either is computed.

Sandboxed execution is a separate package. A Sandbox is configured with a timeout in milliseconds, maxMemoryMB, allowedDomains, stepLimit and captureOutput, then run with a task and a model. The result exposes metrics covering steps, URLs visited and CPU time. That is the piece aimed at unattended runs: a memory ceiling and a domain allowlist bound what a misbehaving agent can reach.

## Installing it and running a first task

The README's Quick Start assumes Bun. Clone the repository, then install at the workspace root, because package.json declares workspaces under packages/* and the root scripts delegate with bun run --filter '*'.

```bash
bun install
```

Next, copy the environment template and fill in at least one provider key. The .env.example lists OPENAI_API_KEY, ANTHROPIC_API_KEY and GOOGLE_AI_API_KEY, and the README's configuration section names GOOGLE_GENERATIVE_AI_API_KEY instead. Those two names do not agree, so check which one the code reads before you set it.

```bash
cp .env.example .env
```

With a key in place, run a task. The default model is gpt-4o and the default step limit is 25 for the CLI, while the library default for stepLimit is 100.

```bash
bun run open-browser run "Find the top story on Hacker News and summarize it"
```

If you would rather drive the browser yourself, the interactive command drops you into a browser> prompt where you can open a URL, extract content as markdown, click a selector and take a screenshot. The README shows a session against news.ycombinator.com using open, extract, click .morelink and screenshot front-page.png. The extract command takes a goal in natural language rather than a CSS selector, which is the part that distinguishes this REPL from a plain Playwright console.

```bash
bun run open-browser interactive
```

For library use, the README's example imports Agent, createViewport and createModel, builds a viewport with headless true, and calls agent.run(). The settings shown are stepLimit 50 and enableScreenshots true.

## Where the project is thin

There are no retrieved releases, so there is no changelog to read and no version to pin. The README says "Production-ready since v1.0", but nothing in the repository metadata corroborates a v1.0 tag. If your dependency policy requires tagged versions, this is a blocker rather than an inconvenience.

The last push was on 2026-04-02. That is more than five months before today, so treat the project as quiet rather than actively developed. The README states contributions are welcome, but a quiet repository with no releases means you should expect to read the source when behaviour surprises you.

Two documented details are incomplete. The README's own example asks the agent to "star it" on GitHub, which requires an authenticated session; the documentation does not describe how authentication state is provided. And relaxedSecurity is described only as disabling browser security features, without saying which ones or what breaks when it is on. The README also does not document rollback, retries at the network level, or how stall detection decides a run is stuck.

Cost is the other honest constraint. Every step sends page context, and screenshots are included by default, so a 100-step run with images is a materially different bill from a 25-step run without them.

## Open Browser versus writing Playwright directly

Playwright is the underlying engine here, so the comparison is not engine against engine. It is whether you write the decision logic or the model does. A Playwright script encodes each step, which makes it fast, deterministic, free to run and easy to debug. It also breaks the moment the page changes, and it cannot handle a flow you did not anticipate.

Open Browser inverts that. You write a task string, and the agent chooses selectors and actions at runtime. That handles variation, and it costs tokens and latency on every run. Playwright's own documentation recommends stable locators and explicit waits; an agent loop replaces that discipline with a model that re-reads the page each step.

The practical split: use Playwright for pages you control or flows you run often enough to maintain, and use Open Browser for one-off extraction, exploratory flows, or sites whose markup you do not want to track. If you already have a Playwright suite, the CLI's browser commands (open, click, type, screenshot, eval, extract, state, sessions) can sit alongside it without a rewrite, since the underlying browser is Playwright.

## Licence and what upgrades cost you

The repository is MIT licensed, with the LICENSE file at the root. That permits commercial use and modification, and it carries no copyleft obligation on your own code. It also means no warranty. This is not legal advice; read the LICENSE file yourself if the distinction matters to your organisation.

The upgrade cost is unusual for an MIT project. With no releases and no changelog, there is no version boundary to upgrade across, so the practical model is to pin a commit hash and move deliberately. The monorepo layout helps here: open-browser, @open-browser/cli and @open-browser/sandbox version independently, so a change in the sandbox package does not force a change in the core agent. The root package.json also declares trustedDependencies for @biomejs/biome, which matters if you run installs in a restricted CI environment where postinstall scripts are blocked by default. Linting and formatting go through Biome, not ESLint or Prettier, so a team with an existing JS toolchain will need to accommodate a second linter or drop Biome from the loop.

## Conclusion

Adopt it if you already run TypeScript and Bun and want an agent loop over Playwright without writing the loop yourself, especially if you need per-run cost tracking or a sandbox with a memory ceiling. Do not adopt it if you need a stable published release, a documented upgrade path, or a supported Python stack. Before committing, verify three things against the repository: that the model identifiers your provider expects match the defaults in the CLI, that the sandbox limits behave as you need on your platform, and that the absence of releases and changelogs is acceptable for your dependency policy.

## FAQ

### What is Open Browser?

It is an MIT-licensed TypeScript framework that gives an AI agent control of a Playwright browser so it can click, type, navigate and extract data to complete a task described in natural language. It ships as a monorepo with a core library, a CLI and a sandbox package.

### How do I install Open Browser?

The README's Quick Start uses Bun: run bun install at the repository root, copy .env.example to .env and add at least one provider API key. Then run a task with bun run open-browser run followed by the task in quotes.

### How do I use Open Browser?

You can run a one-off task from the CLI, drop into the interactive REPL with open-browser interactive, or import Agent, createViewport and createModel and call agent.run() from your own TypeScript. The CLI also exposes browser commands such as open, click, type, screenshot, eval and extract.

### How does Open Browser compare with Playwright?

Open Browser is built on Playwright, so the difference is who chooses the steps. A Playwright script encodes each action and is deterministic, while Open Browser sends page context to a model that picks actions at runtime, which handles unfamiliar pages at the cost of tokens and latency per run.

## Sources

- [Issues](https://github.com/ntegrals/openbrowser/issues)
- [License: MIT](https://github.com/ntegrals/openbrowser/blob/master/LICENSE)
- [ntegrals/openbrowser on GitHub](https://github.com/ntegrals/openbrowser)
- [README](https://github.com/ntegrals/openbrowser/blob/master/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ntegrals-openbrowser
