# browser-use-mcp-server: browser automation for Cursor, Windsurf and Claude

> An MIT-licensed MCP server that hands browser-use agents to any MCP client, with SSE and stdio transports, a Docker image with VNC, and an OpenAI key requirement that shapes who can actually use it.

**kontext-security/browser-use-mcp-server** — Browse the web, directly from Cursor etc.

- Repository: https://github.com/kontext-security/browser-use-mcp-server
- Website: https://kontext.security
- Stars: 845 · Forks: 116
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/kontext-security-browser-use-mcp-server

## The gap browser-use-mcp-server fills between an editor and a browser

Most coding assistants can read files and run shell commands. They cannot click a link, wait for a page to render, or read what a JavaScript-heavy site produced after hydration. browser-use-mcp-server closes that gap by exposing browser-use as a Model Context Protocol server, so an MCP client such as Cursor, Windsurf or Claude Desktop can drive a real Chromium instance.

The target user is narrow and identifiable: someone already running an MCP client who wants page interaction inside the same session as their code. The README's canonical example is a one-line request to the assistant, asking it to open a URL and return the top ranked article. That is the shape of the work it is built for: fetch, read, extract, report. It is not a scraping framework and it is not a test runner. The repository description puts it plainly as browsing the web from Cursor and similar tools.

The project is written in Python and licensed MIT, and the last push to the default branch was on 2026-05-20. It is not archived. Releases are sparse: v1.0.3 landed on 2025-04-15, with v1.0.2 and v1.0.1 earlier the same day. The README still points at co-browser URLs and a cobrowser.xyz support address while the repository now lives under kontext-security, which is worth knowing when you go looking for the issue tracker.

## How the browser-use MCP server moves a request to a page

The architecture is a thin MCP wrapper around browser-use, which itself drives Playwright. The pyproject dependencies make the chain explicit: browser-use>=0.1.40, playwright>=1.50.0, mcp>=1.3.0, plus langchain-openai>=0.3.1 for the model side. Starlette and uvicorn serve the HTTP transport.

The server supports two transports. In SSE mode it listens on a port and the client connects to an /sse URL. In stdio mode the client launches the process itself and speaks over standard input and output, with mcp-proxy listed as a prerequisite for that mode. Two ports appear in the documented commands: 8000 for the server and 9000 as the proxy port in stdio mode.

Model calls go to OpenAI. The .env.example defines exactly four keys: CHROME_PATH, OPENAI_API_KEY, PATIENT and ANONYMIZED_TELEMETRY. PATIENT is the interesting one. The comment in the example file says it controls whether API calls wait for task completion, defaulting to false, which means the default behaviour is asynchronous task execution rather than a blocking call. The README lists async tasks as a feature, and this flag is how you opt out of it.

The Dockerfile adds a second surface: a runtime image built on debian:bookworm-slim with xfce4, tigervnc-standalone-server and proxy-login-automator, exposing a VNC port so you can watch the automation. The image sets ANONYMIZED_TELEMETRY=false in its environment, and creates a fallback VNC password file at /run/secrets/vnc_password_default containing browser-use.

## Installing browser-use-mcp-server with uv and running a first task

The README's prerequisites are uv, Playwright and mcp-proxy. It gives install commands for all three, including a shell installer for uv and a global tool install for the proxy.

```bash
curl -LsSf https://astral.sh/uv/install.sh | sh
uv tool install mcp-proxy
uv tool update-shell
```

After that, create a .env file. Only OPENAI_API_KEY is strictly needed for a first run; the other three keys have defaults or are optional. CHROME_PATH is left blank to use the Playwright-managed Chromium.

```bash
OPENAI_API_KEY=your-api-key
CHROME_PATH=
PATIENT=false
ANONYMIZED_TELEMETRY=false
```

Then sync dependencies and install the browser binary. The --no-shell flag in the README's command skips the Playwright shell, which matters on machines where you only need Chromium.

```bash
uv sync
uv pip install playwright
uv run playwright install --with-deps --no-shell chromium
```

Start the server in SSE mode on port 8000. The README shows the source-run form.

```bash
uv run server --port 8000
```

Point your client at it. For Cursor the README gives ./.cursor/mcp.json as the configuration path, and the SSE entry is a single URL.

```json
{
  "mcpServers": {
    "browser-use-mcp-server": {
      "url": "http://localhost:8000/sse"
    }
  }
}
```

If you prefer stdio, build a wheel, install it as a global tool, and run it with the stdio and proxy-port flags. The README's exact sequence is below.

```bash
uv build
uv tool uninstall browser-use-mcp-server 2>/dev/null || true
uv tool install dist/browser_use_mcp_server-*.whl
browser-use-mcp-server run server --port 8000 --stdio --proxy-port 9000
```

The first task to try is the README's own example: ask the assistant to open https://news.ycombinator.com and return the top ranked article. With PATIENT=false the call returns before the page work finishes, so expect to poll or wait for the task rather than get an answer in one turn.

## Docker and the VNC view of a running browser

The Docker path is the more reproducible one, and it is the only place the project documents a way to watch what the agent is doing. The README builds the image and runs it with two port mappings: 8000 for the server and 5900 for VNC.

```bash
docker build -t browser-use-mcp-server .
docker run --rm -p8000:8000 -p5900:5900 browser-use-mcp-server
```

To override the password, write it to a file and mount it read-only at /run/secrets/vnc_password. The README notes that the :ro flag makes the file read-only inside the container.

```bash
echo "your-secure-password" > vnc_password.txt
docker run --rm -p8000:8000 -p5900:5900 \
  -v $(pwd)/vnc_password.txt:/run/secrets/vnc_password:ro \
  browser-use-mcp-server
```

For a viewer, the README clones noVNC and runs its proxy against localhost:5900. The default password is browser-use unless you used the secrets mount. This is the part of the project with the clearest operational value: when an agent-driven browser gets stuck on a consent dialog or a login wall, a VNC session tells you that in seconds, and a log file usually does not. Note that the image installs a full xfce4 desktop and a VNC server, so the runtime image is far heavier than a headless Playwright container would be.

## Where browser-use-mcp-server is the wrong tool

The OpenAI dependency is the first hard boundary. OPENAI_API_KEY is the only model credential the configuration documents, and langchain-openai is a direct dependency. If your constraint is a local model or a non-OpenAI provider, nothing in the README or the .env.example tells you how to point the agent elsewhere. That is not a small gap: it decides whether the project fits a regulated environment at all.

Second, the async default. With PATIENT=false the server returns before the task completes, and the README does not describe how a client is expected to collect the result. If your workflow assumes one request maps to one finished answer, you are fighting the default rather than using it.

Third, platform coverage. The README lists configuration paths for Cursor, Windsurf, Claude on macOS and Claude on Windows, but every install and run command is a POSIX shell command with uv, and the Docker image is Debian-based. There is no Windows-native setup path documented.

Fourth, this is not a scraper. It drives a full browser through an LLM, which means each task costs model tokens and wall-clock time. For a fixed extraction job against a known DOM, a plain HTTP client or a Playwright script is cheaper, faster and far easier to make deterministic. Reach for this when the page is unknown in advance or the steps depend on what the assistant reads.

Finally, the README does not document rollback, version pinning beyond the wheel glob, or a migration path between releases. The release history is thin, so treat upgrades as something you verify on a branch rather than assume.

## How it differs from a browser extension MCP bridge

The obvious alternative, and one the related searches keep circling, is a Chrome extension that exposes MCP to your existing browser session. The difference in approach is fundamental. An extension bridge attaches to the browser you already use, so it inherits your cookies, your logged-in sessions and your extensions. That is convenient and also means the agent acts as you, inside a profile full of credentials.

browser-use-mcp-server goes the other way. It launches its own Chromium through Playwright, in its own profile, and in the Docker case inside its own container. Sessions do not leak in from your daily browser, and nothing the agent does touches your personal logins unless you deliberately authenticate inside the agent's browser. The cost is that any site requiring a login must be logged into again inside that controlled browser, and the VNC view exists precisely because you may need to do that by hand.

A second alternative is writing Playwright scripts directly. Those are deterministic, cheap to run and easy to review. They also cannot decide mid-run that a page looks different today and adapt. The trade you are making with browser-use-mcp-server is model cost and nondeterminism in exchange for handling pages you did not anticipate. If you cannot name a case where the page changes under you, you probably do not need this.

## Maintenance, licence and what an upgrade actually costs

The licence is MIT, declared in pyproject.toml as license = {text = "MIT"} and shipped as a LICENSE file at the repository root. MIT is permissive: you can use, modify and redistribute the code, including commercially, provided the copyright notice and permission notice travel with it. That is a statement about the licence text, not legal advice; if you are redistributing the Docker image or a modified wheel, have your own counsel read the notice requirements.

The dependency chain is where upgrade cost actually lives. browser-use, playwright, mcp and langchain-openai are all fast-moving, and the project pins only lower bounds (browser-use>=0.1.40, playwright>=1.50.0, mcp>=1.3.0). A fresh uv sync months from now can resolve to versions the code was never exercised against. The repository does carry a uv.lock, and the Dockerfile builds with uv sync --frozen, so the container path is reproducible even when a local sync is not. If you care about repeatable installs, prefer the frozen Docker build over a bare uv sync.

On maintenance signals: the last push was on 2026-05-20, and the most recent tagged release is v1.0.3 from 2025-04-15. The gap between the two is the honest picture. The repository is not archived, but the release cadence is not something you can plan around, and the README still references co-browser URLs and a cobrowser.xyz support address while the project sits under kontext-security. Budget for reading the source when something breaks, because the documentation will not always match the code you pull.

## Conclusion

Adopt it if you already run an MCP client, want browser tasks inside that client, and are willing to supply an OPENAI_API_KEY and a Chromium install. Skip it if you need a provider-agnostic model layer, a Windows-native path, or a documented rollback story; the README covers none of those. Before wiring it into a team config, verify three things yourself: that uv sync and uv run playwright install --with-deps --no-shell chromium complete on your machine, that port 8000 (or your chosen port) is free, and that the stdio path works after uv build and uv tool install dist/browser_use_mcp_server-*.whl.

## FAQ

### Can I use MCP to control my browser with browser-use-mcp-server?

Yes. The project is an MCP server that drives a browser through browser-use, and the README documents both an SSE mode on port 8000 and a stdio mode that the client launches directly. In SSE mode the client connects to http://localhost:8000/sse.

### Where can I run the browser-use-mcp-server?

The README documents three places: from source with uv run server --port 8000, as an installed global tool via uv build and uv tool install dist/browser_use_mcp_server-*.whl, or in Docker with docker run --rm -p8000:8000 -p5900:5900 browser-use-mcp-server. The Docker image is Debian-based and also exposes VNC on port 5900.

### Can I use browser-use-mcp-server with a local LLM?

OPENAI_API_KEY is the only model credential in .env.example, and langchain-openai is a direct dependency, with no documented way to point the agent at a different provider. The README does not describe a local model path.

## Sources

- [kontext-security/browser-use-mcp-server on GitHub](https://github.com/kontext-security/browser-use-mcp-server)
- [License: MIT](https://github.com/kontext-security/browser-use-mcp-server/blob/main/LICENSE)
- [Project website](https://kontext.security)
- [README](https://github.com/kontext-security/browser-use-mcp-server/blob/main/README.md)
- [Releases](https://github.com/kontext-security/browser-use-mcp-server/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/kontext-security-browser-use-mcp-server
