Open-source project
web-infra-dev/midscene avatar
web-infra-dev/midscene

Midscene.js: vision-driven UI automation for Playwright, Vitest and AI agents

AI-powered, vision-driven UI automation for every platform. Two ways to test, add Midscene to your Playwright / Vitest suite, or let an AI agent test autonomously via Skills.

15,016 stars1,170 forksTypeScriptMIT

At a glance

What is it?
Midscene.js targets elements from screenshots instead of selectors, and runs either inside a Playwright or Vitest suite or as an autonomous agent via Skills. Here is how the mechanism works, what it costs, and where it stops being the right tool.
Who is it for?
Adopt Midscene if your UI resists selectors (canvas, icon-only controls, native apps, cross-origin iframes) or if you want assertions about what a screen actually looks like, and accept that every run depends on a multimodal model you must configure and pay for. Do not adopt it as a drop-in replacement for a selector-based suite in a project with no model budget, no screenshot pipeline, or a hard requirement that tests be deterministic and offline.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 6 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The selector problem Midscene.js is built around

Most UI automation, including AI tools that read the DOM or the accessibility tree, depends on page structure. The README argues that this structure is fragile and incomplete: selectors break on refactors, elements without semantic markup are invisible to it, and native apps and cross-origin iframes are out of reach. The stated consequence is that such tools cannot tell whether something actually looks right. Midscene takes the opposite position: element localization is based on screenshots only, and each step is described in natural language. That makes the intended audience fairly narrow and fairly specific. It is for teams already writing end-to-end tests who keep repairing selectors, and for teams testing surfaces where selectors were never available in the first place: canvas rendering, icon-only buttons, custom controls, Android, iOS, HarmonyOS and desktop apps. If your test suite is stable and your UI exposes clean test ids, the pitch does not apply to you.

Screenshot in, localized element out: the actual mechanism

The README is explicit that Midscene is all-in on pure vision for UI actions. The flow is: capture a screenshot of the current surface, send it to a multimodal model with your natural-language instruction, receive back a localization of the target element, then act on it. Midscene runs on models with strong UI localization, and the README names Qwen3.x, Doubao-Seed-2.1, GLM-4.6V, gemini-3.5-flash and UI-TARS, including open-source options you can self-host. There is one documented escape hatch: for data extraction and page understanding you can still opt in to include DOM when needed. So the DOM is not banned outright, it is demoted from the source of truth for actions to an optional input for reading. The API surface reflects this split. The README points at aiAct, aiQuery and aiAssert as the common methods, which map to acting on the screen, extracting structured data, and asserting what is visible. Assertions are where the vision approach pays off most clearly: verifying colors, highlights, layout and rendered state is not something a DOM-node-exists check can express. The cost is equally clear. Every step is a model call against a screenshot, so latency and spend scale with the number of steps, and the quality of localization is a property of the model you choose, not of the library.

Getting started with Midscene in a Playwright suite

The README does not print install commands; it routes you to the quick start, the Playwright and Puppeteer integration guides, and the platform guides for Android, iOS, HarmonyOS and desktop. The npm package it links is @midscene/web, and the repository is a pnpm workspace, so confirm the current install line on the documentation site before you copy anything. The quick start describes configuring a model, installing the Chrome extension, and running a first natural-language instruction, which is the cheapest way to find out whether your chosen model can localize your UI at all. That is also the only piece of setup the README spells out end to end, so a reader who wants a script instead of the extension has to follow the Playwright guide on the site. The repository's own scripts show how the project is driven during development: the root package.json defines a build target over the workspace and a dev target that watches all packages except android-playground, chrome-extension, @midscene/report and doc.

bash
npm run build
npm run dev

Those commands build and watch the monorepo itself rather than a consumer project, so they are useful when you want to read or patch the source, not when you are adding Midscene to an existing suite. What you should expect after a consumer-side install is an Agent object that drives the page from screenshots and returns a pass or fail for an assertion. The constructor options and the model configuration keys live in the API reference rather than the README, so check the reference for the fields your model provider needs.

Two ways to run it, and why the Skills path changes the economics

Midscene offers two modes that are easy to conflate. The first is as a library inside a suite you already own: add it to Playwright or Vitest, and your test runner keeps control of ordering, retries and reporting. The second is Midscene Skills, which the README describes as letting an AI agent test autonomously, used with OpenClaw to test and automate web, mobile and desktop interfaces. The difference matters more than it looks. In the library mode, a human writes the steps and the model only localizes and judges; the test suite stays the artifact of record. In the Skills mode, the agent decides what to do, which is useful for exploratory passes and much harder to pin to a fixed regression baseline. The README also states that the same vision-driven engine handles any UI automation task, not just testing, and that you can write automation with the JavaScript SDK or in YAML. YAML is the lower-friction option for people who do not want to write TypeScript; the SDK is the option when you need to branch on results inside a larger program.

Where Midscene is the wrong tool

The README does not document rollback, retry semantics, or how a failed localization is reported to the test runner, and that silence is worth taking seriously before you put it on a release gate. The larger constraint is structural: because actions are localized from screenshots by a model, a run is only as reproducible as the model behind it. A hosted model can change under you, and a self-hosted one needs the hardware and the operational attention that the README does not cover. Two failure modes follow. First, cost and latency grow with step count, so a long flow that a selector-based test finishes in seconds becomes a sequence of model calls. Second, screenshot-driven localization has no deterministic fallback: when the model misreads a dense screen, you get a wrong click or a failed assertion with no selector to inspect. Teams with strict offline or air-gapped CI, teams whose UI is already well annotated, and teams that need byte-identical test runs should stay with Playwright's own locators. Midscene is a complement to that suite, not a replacement for it.

Midscene.js compared with DOM-reading agents

The closest alternatives are agents that drive the browser through the DOM or the accessibility tree, the category Browser Use and Stagehand belong to. The difference in approach is where the element comes from. A DOM-reading agent parses structured markup and picks a node, which is fast, cheap and inspectable, and which collapses the moment the target has no semantic markup or lives in a cross-origin iframe. Midscene reads the rendered pixels instead, so the same instruction works on a canvas element, a native Android screen and a web page without three different strategies, and assertions can cover appearance rather than existence. The trade is the one described above: you pay per step in model calls, and correctness depends on the model's localization ability. For data extraction specifically, Midscene's documented opt-in to include DOM narrows the gap, because reading structured content is exactly what the DOM is good at. The README also lists community ports, including a Python SDK, two Java SDKs, an iOS mirror project and a Docker image with a PC server, so the ecosystem extends beyond the TypeScript core even though the repository itself is TypeScript.

Maintenance, licensing and what upgrading costs

The repository is not archived and the last push was on 2026-08-28, the same day as the v1.12.2 release, with v1.12.1 and v1.12.0 landing earlier that month. That is a recent release cadence, and the root package.json carries version 1.12.6, which suggests the workspace moves ahead of the published tags. The project is MIT licensed, which permits commercial use and modification; that is a statement about the licence text, not legal advice, and if you redistribute it inside a product you should read the LICENSE file yourself. The upgrade cost is dominated by the model layer rather than the library. Midscene names specific models, and model availability, pricing and naming change on the provider's schedule, not Midscene's, so a version bump can quietly change localization quality. The repository is an Nx monorepo with a pnpm workspace, which means contributing or building from source pulls in the full toolchain; consuming it as @midscene/web does not. One practical detail for anyone reading the source: the root package.json defines test:ai targets scoped to @midscene/core, @midscene/web and @midscene/cli, so the project's own AI-driven tests are separable from its unit tests, and running them will require model credentials.

Editorial conclusion

Adopt Midscene if your UI resists selectors (canvas, icon-only controls, native apps, cross-origin iframes) or if you want assertions about what a screen actually looks like, and accept that every run depends on a multimodal model you must configure and pay for. Do not adopt it as a drop-in replacement for a selector-based suite in a project with no model budget, no screenshot pipeline, or a hard requirement that tests be deterministic and offline. Before committing, verify which model your team can actually reach, whether the YAML path or the JavaScript SDK fits your repository, and whether the AI-driven test targets in package.json are gated behind credentials you do not have.

Frequently asked questions

What is Midscene.js?

Midscene.js is an AI-powered, vision-driven UI automation tool for end-to-end testing. It localizes elements from screenshots rather than selectors, and it can be added to a Playwright or Vitest suite or run autonomously through Midscene Skills.

How does Midscene.js differ from Browser Use?

Browser Use belongs to the category of agents that read the DOM or accessibility tree, while Midscene states that element localization is based on screenshots only. Midscene does allow opting in to include DOM for data extraction and page understanding.

How does Midscene.js compare with Stagehand?

The README groups DOM-reading AI tools together as depending on page structure, which is the approach Midscene positions itself against. Midscene works from the screenshot alone, so it can target canvas elements, native apps and cross-origin iframes that a structure-based agent cannot reach.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/web-infra-dev-midscene.svg)](https://hysenlabs.com/projects/web-infra-dev-midscene)