Model or dataset
AIPexStudio/AIPex avatar
AIPexStudio/AIPex

AIPex automates the browser you already have, over a local websocket on port 9223

AIPex: AI browser automation assistant, no migration and privacy first. Alternative to Manus Browser Operator、 Claude Chrome and Agent Browser

1,252 stars132 forksTypeScriptMIT

At a glance

What is it?
An MIT licensed browser extension that gives coding agents and scripts control of your existing browser instead of a separate automation profile. Agents reach it through a bridge process over stdio and a websocket on port 9223, pages are handled as text snapshots with stable element ids rather than screenshots, and the repository carries nine translations, a skill package and a version number that stopped at 0.0.2.
Who is it for?
AIPex is worth a look if the thing stopping you from browser agents is that they want a second browser, a subscription or your history. The architecture is coherent: one local daemon, a bridge over stdio, structured snapshots instead of image loops, and state that never enters your repository.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 38 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The argument is that your browser already works, so why install another one

AIPex is a browser extension rather than a browser, and the readme frames that as the whole idea.

The claim is zero migration: no new browser to install and no new workflow to learn. Three costs are named as belonging to the alternatives. Separate browsers such as Dia and Comet have to be installed and switched into. Subscription products such as ChatGPT Atlas are paid monthly. And the third cost is your browsing data, which the project treats as something you should not have to hand over to get automation.

The response is to run the agent inside the profile you already use, where your logged in sessions, cookies, tabs, history and other extensions already exist. Alongside that sit two other commitments: an MIT licence so the extension is auditable and extensible, and a privacy position stated as your data never leaves your machine with your own API key.

One detail is worth noticing because it is odd. The repository description names Manus Browser Operator, Claude Chrome and Agent Browser as the alternatives it positions against, and the declared homepage is a domain named after one of them. Whether that is a redirect, a shared owner or a leftover is not explained anywhere in the repository.

An agent talks to a bridge over stdio, and the bridge talks to the extension over a websocket

The control path is four hops, and every one of them is local.

An AI agent starts a process called the bridge, communicates with it over stdio, and the bridge opens a websocket to the extension, which then drives the browser through its own APIs. Because the agent side is stdio, the configuration is a command line, which is why the same bridge works in editors that speak different configuration formats.

Three of them share one JSON shape. Cursor uses a project MCP file, Claude Desktop uses its desktop config file, and Windsurf uses its own config file; all three declare a server named for the browser, with the command set to npx and the arguments pointing at the bridge package with its yes flag. Claude Code takes the same thing as a single command that adds the server and then runs the package. VS Code Copilot is configured in a project file too, but under a different key, so its snippet is not a copy of the other three.

Then the extension side. Open the extension options, set the websocket URL to the local extension path on port 9223, and connect. After that the agent has more than thirty browser tools available.

The port is fixed and the transport is an unencrypted local websocket, which is a deliberate simplification and also the reason this only works on your own machine.

Page state is a text snapshot you can grep, not an image you have to look at

The performance argument is about what the agent has to read, and it is the most interesting engineering decision in the project.

Before taking a screenshot, AIPex produces a structured snapshot of the page. Agents search those snapshots with glob and grep patterns and act on elements by stable identifiers rather than by position. So finding every button on a page is a pattern match against text, and clicking one is a call naming an id that survives a re-render.

The stated payoff is fewer slow screenshot loops for most tasks, and lower token and latency overhead, because sending text and targeting one element is much cheaper than repeatedly shipping full page images to a model. Three items on the roadmap confirm the direction of travel: an optimised DOM, dropping unused snapshot content, and id based operation.

What is not done is vision. It is the one unchecked box under page understanding, sitting alongside a completed accessibility tree. So the project is explicitly structured-first: a page that renders its content as DOM rather than canvas is in scope, and anything that only exists as pixels is not, because nothing in the current design looks at pixels.

That is also the failure mode to keep in mind, since a canvas based application defeats the snapshot before any model gets involved.

browser-cli was a separate tool and is now part of the same daemon

There are two command line surfaces, and the friendly one arrived by absorption.

browser-cli used to be its own thing and has been merged into AIPex. It now ships inside the bridge package and talks to the same local daemon that MCP uses, which is the point of the merge: scripts, continuous integration jobs and coding agents all drive one browser runtime instead of standing up separate services.

bash
npm install -g aipex-mcp-bridge

browser-cli status
browser-cli tab list
browser-cli tab new https://example.com
browser-cli page search "button*" --tab 123
browser-cli interact click btn-42 --tab 123

The shape of those commands is the design. A status check, then tab listing and creation, then a page search with a glob and a tab number, then an interaction that names an element id. Every command carries the tab it applies to, because the browser you are automating has many.

Underneath, the groups are tab, page, interact, download, intervention and skill. Intervention is the one to look at, since it implies a way to hand control back to a human mid task. A lower level command line still exists for raw tool calls, so the friendly layer is a convenience rather than a replacement.

Installing from npm runs a git hook installer, and the autofix flag is marked unsafe

The repository is a pnpm workspace with a private root package, and its scripts say more about the working style than a contributing guide would.

The root package carries no runtime dependencies at all, only development ones, which is what a workspace root looks like when the real packages live under the packages directory. The version in that root is 0.0.2 while the newest published release is 0.1.0, so the root version is not the product version.

Two scripts deserve attention. The postinstall hook runs a pre-commit installer, so a plain install of the workspace writes git hooks into your checkout without asking. And the lint fix script passes an unsafe flag to the checker, which means the automated fixes it applies are the ones the tool itself classifies as behaviour changing. Both are ordinary choices in a monorepo, and both are the kind of thing you want to know about before you run an install in a repository you care about.

The quality gate itself is a single preflight chain that formats, fixes, typechecks and tests in order, backed by one linter and formatter for both jobs, a dependency checker running in strict mode, and one test runner. Type definitions for the browser and for node are pinned to exact versions in the root.

The contributing section points at a development document, and no such file exists at the repository root.

Nine readmes, one skill package, and no release since March

The project's shape is visible in its file list, and it is a consumer extension first and a library second.

Nine readme files sit at the root, one per language, which is an unusual amount of translation surface to keep synchronised for an extension whose configuration is a JSON block. Alongside them are two agent instruction files and an editor configuration, so the repository is set up to be worked on by coding agents as well as humans.

There is a skill package as well. It targets runtimes that support the skill protocol, and it bundles tool usage strategy, complete parameter schemas for all thirty plus browser tools, and common automation patterns. The intent is stated plainly: an agent should not have to discover thirty tools by trial and error. That is a real cost when a model guesses argument names, and packaging the schemas is a direct answer to it.

The release history is thin for a project that presents itself as an alternative to shipping browser products. Two patch releases landed two days apart in January, then a 0.1.0 in March, and nothing since, although the last push to main came at the end of August.

So the code is moving and the published version is not, which is worth remembering before you file an issue about a missing fix.

The roadmap has an empty section, two typos, and an integration it forgot to tick

The roadmap is where the project's rough edges are most visible, and three of them are documentation rather than code.

First, there is a section heading that asks why a debugger is necessary for browser automation, and nothing follows it. The section is empty in the readme, which means either the explanation was never written or it was cut.

Second, the completed boxes contain misspellings, in the accessibility tree entry and in the search based retrieval entry. Small, but they are the kind of thing that survives review when nobody reads the words rather than the boxes.

Third, the integration list has two ticked entries, Cursor and Claude Code, while the configuration section documents four hosts including Claude Desktop, Windsurf and VS Code Copilot. So the shipped support and the roadmap disagree about what is finished.

The unchecked items are more interesting. Vision is not done, and neither is evaluation, which is marked against a public web agent benchmark rather than a private test set. Choosing an external benchmark before shipping one is a defensible order of work, and it also means there is no published accuracy claim anywhere in the repository to check the demos against.

Editorial conclusion

AIPex is worth a look if the thing stopping you from browser agents is that they want a second browser, a subscription or your history. The architecture is coherent: one local daemon, a bridge over stdio, structured snapshots instead of image loops, and state that never enters your repository. Two things to check before you rely on it. Vision is still unchecked on the roadmap, so anything that genuinely needs to read a rendered page rather than a DOM tree is outside what the project claims to do. And the version numbers disagree, with the root package sitting at 0.0.2 while the newest published release is 0.1.0, so pin the extension from the store and read the release notes rather than assuming the two match.

Frequently asked questions

What is AIPex and how is it different from other browser agents?

AIPex is an open source browser automation agent that runs inside your existing browser rather than a separate automation profile, so there is no new browser to install and your logged in sessions, cookies and tabs are already there. It is MIT licensed and states a privacy first position: your data stays on your machine and you bring your own API key.

How do I install AIPex?

Install the extension from the Chrome Web Store or Edge Add-ons, press the extension icon, then type or speak what you want in natural language. For the command line and MCP bridge, install the package globally with npm, which provides both the bridge process and the browser-cli command.

How do coding agents control AIPex through MCP?

Add the bridge as an MCP server over stdio. Cursor, Claude Desktop and Windsurf use a JSON config with a servers entry whose command is npx and whose arguments run the bridge package; Claude Code takes the same setup as a single command. Then open the extension options, set the websocket URL to the local extension path on port 9223, and connect, after which more than thirty browser tools are exposed.

What is the difference between browser-cli and aipex-cli?

browser-cli was merged into AIPex and now ships with the bridge package, talking to the same local daemon, with human friendly command groups for tabs, pages, interaction, downloads, intervention and skills. The lower level aipex-cli remains available for raw tool calls, so the friendly layer sits on top rather than replacing it.

Does AIPex support agents that use a skill protocol?

Yes. It ships an aipex-browser skill package for runtimes that support the skill protocol, and it bundles tool usage strategy, complete parameter schemas for all thirty plus browser tools, and common automation patterns, so an agent does not have to discover the tool surface from scratch.

Does AIPex work with vision or only with page structure?

Only with structure, for now. The design takes a DOM snapshot before resorting to images, and agents search those snapshots with glob and grep patterns and act on stable element identifiers. Vision is the one unchecked item under page understanding on the roadmap, so pages that render their content as pixels rather than DOM are outside what the project claims to handle.

Official sources

  1. AIPexStudio/AIPex on GitHub
  2. License: MIT
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/aipexstudio-aipex.svg)](https://hysenlabs.com/projects/aipexstudio-aipex)