Page Agent: A JavaScript In-Page GUI Agent for Web Automation Without Browser Extensions
JavaScript in-page GUI agent. Control web interfaces with natural language.
At a glance
- What is it?
- Page Agent is an open-source TypeScript library from Alibaba that embeds an AI agent directly into a webpage using plain JavaScript. It controls web interfaces through text-based DOM manipulation without screenshots, headless browsers, or Python, making it a practical option for adding natural-language automation to existing web apps.
- Who is it for?
- Page Agent suits frontend developers who need to add AI-driven automation to a web application without changing the server or installing a browser extension. The library works client-side, so it integrates with any existing web stack through a script tag or npm package.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
The Problem Page Agent Solves
Most web automation tools operate from outside the browser. They require a Python runtime, a headless browser such as Playwright or Puppeteer, or a browser extension installed on the user's machine. Each of these adds setup friction and is difficult to ship as part of a hosted web product.
Page Agent takes a different position: all automation runs inside the existing webpage as plain JavaScript. There is no extension to install, no Python dependency, no headless browser process. According to the README, everything happens in the web page itself.
This design makes Page Agent suited for teams that want to ship an AI copilot feature in a SaaS product without adding infrastructure. The README names smart form filling for ERP, CRM, and admin systems as a target use case, alongside accessibility applications where natural language commands replace multi-click workflows.
Text-Based DOM Manipulation Instead of Screenshots
Many in-browser AI agents work by taking a screenshot, sending it to a multimodal model, and then interpreting the model's instructions as DOM actions. Page Agent does not use screenshots. It processes the DOM as text.
The README describes this as text-based DOM manipulation with no screenshots and no multi-modal LLMs or special permissions needed. This has a direct practical effect: it makes the library compatible with models that have no vision capability, and it avoids the latency of capturing and uploading a screenshot on every action step.
The DOM processing components and prompt design are derived from the browser-use project (github.com/browser-use/browser-use), which the README acknowledges explicitly. The README states: "DOM processing components and prompt are derived from browser-use: Browser Use <https://github.com/browser-use/browser-use>, Copyright (c) 2024 Gregor Zunic, Licensed under the MIT License."
Installing and Running a First Automation
For evaluation, the fastest option is a single script tag that loads the demo build from a CDN:
<script
src="https://cdn.jsdelivr.net/npm/[email protected]/dist/iife/page-agent.demo.js"
crossorigin="anonymous"
></script>The README warns that this demo CDN uses a free testing LLM API intended for technical evaluation only, not production use.
For production use, install via npm:
npm install page-agentThen import and instantiate the agent with your own LLM credentials:
import { PageAgent } from 'page-agent'
const agent = new PageAgent({
model: 'qwen3.5-plus',
baseURL: 'https://dashscope.aliyuncs.com/compatible-mode/v1',
apiKey: 'YOUR_API_KEY',
language: 'en-US',
})
await agent.execute('Click the login button')The constructor takes `model`, `baseURL`, `apiKey`, and `language`. The `baseURL` accepts any OpenAI-compatible endpoint, which is how it supports locally deployed models alongside cloud providers.
Multi-Page Tasks via the Chrome Extension and MCP Server
The base library handles automation within a single page. For tasks that span multiple browser tabs or require coordination across different URLs, Page Agent provides an optional Chrome extension listed in the Chrome Web Store with extension ID `akldabonmimlicnjlflnapfeklbfemhj`.
The README also mentions an MCP Server in beta that allows agent clients to control the browser from outside the page. This is a different integration model from the in-page library: where the core library runs inside the webpage, the MCP Server exposes browser control as a tool that an external agent can call.
The monorepo is organized as a set of npm workspaces. The root `package.json` lists these packages: `page-controller`, `ui`, `llms`, `core`, `page-agent`, `mcp`, `extension`, and `website`. The separation between `core` and `page-agent` suggests that the DOM processing logic is in a shared package consumed by both the main library and the extension.
Supported Models and the Bring-Your-Own-LLM Constraint
Page Agent works with most mainstream models, including locally deployed ones, according to the README. The documentation site lists supported models at `alibaba.github.io/page-agent/docs/features/models`.
The constructor's `baseURL` parameter accepts any OpenAI-compatible endpoint. The example in the README uses `dashscope.aliyuncs.com/compatible-mode/v1` with `qwen3.5-plus`, which is a Qwen model served through Alibaba Cloud's DashScope API. A developer using a locally deployed model through Ollama or LM Studio could point `baseURL` at their local server if that server exposes an OpenAI-compatible API.
The free demo CDN uses a testing LLM API with its own terms, documented at the Page Agent terms-and-privacy page. The README's warning is clear: it is for technical evaluation only. Any production deployment requires the developer to supply credentials for a supported model.
Limitations and Cases Where Page Agent Is the Wrong Tool
Page Agent is explicitly not designed for server-side automation. The README states: "PageAgent is designed for client-side web enhancement, not server-side automation." If you need to automate a website from a server process, tools like Playwright or browser-use are the appropriate choice.
The library also has no built-in handling for pages that use heavy canvas rendering, PDF viewers, or iframe-heavy layouts where text-based DOM access is limited. The README does not document how the agent behaves on these page types.
Contributions generated entirely by bots or AI without substantial human involvement are not accepted, according to the README. This is a stated policy constraint for the project, not a technical limitation, but it is relevant for teams evaluating whether they can submit automated patches.
The project requires Node.js 22.22.1 or later, or Node.js 24 or later. The root `package.json` sets `engines: { node: "^22.22.1 || >=24" }`. Environments running older Node.js versions will need to upgrade before using the library.
License, Maintenance, and Project State
Page Agent is licensed under MIT. The DOM processing components derived from browser-use are also MIT-licensed, as noted in the README attribution block. There are no license compatibility issues for commercial use of either the library or its browser-use-derived components.
The most recent release is v1.12.4, published on 2026-09-06. The last push to the repository was on 2026-09-21. The project is not archived. Earlier releases in this series include v1.12.3 from 2026-09-05 and v1.12.2 from 2026-07-16, indicating a consistent release cadence in mid-2026.
The repository uses Husky for git hooks and commitlint for commit message validation. The CI pipeline is defined in `.github/workflows/main-ci.yml`. Testing is done through npm workspaces, with each package running its own test suite.
Editorial conclusion
Page Agent suits frontend developers who need to add AI-driven automation to a web application without changing the server or installing a browser extension. The library works client-side, so it integrates with any existing web stack through a script tag or npm package. It is not suited for server-side automation pipelines, which is a stated design constraint: the README explicitly notes that PageAgent is designed for client-side web enhancement, not server-side automation. Before adopting it, check the supported models list at the documentation site to confirm your LLM is compatible. The latest release is v1.12.4 from 2026-09-06 and the last push was on 2026-09-21, so the project is under active development.
Frequently asked questions
What is Page Agent used for?
Page Agent is used to add natural-language automation to existing web applications without a backend rewrite. Common cases include AI copilots in SaaS products, smart form filling in ERP and CRM systems, and accessibility features where voice commands or plain text replace multi-click workflows.
Does Page Agent require a browser extension?
The core library requires no browser extension. Everything runs as in-page JavaScript. An optional Chrome extension is available for tasks that span multiple browser tabs, but single-page automation works without installing anything in the browser.
Which LLMs does Page Agent support?
Page Agent works with most mainstream models, including locally deployed ones. The constructor accepts any OpenAI-compatible baseURL, so it can connect to cloud providers or a local server. The full supported-model list is at the documentation site.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/alibaba-page-agent)
Community notes