# MAI-UI and Qwen-UI-Agent: benchmark numbers and a mobile device farm

> A research repository with almost no code in it, where the contribution is a GUI agent, a benchmark built on 100 physical phones, and a list of scores to read carefully.

**Tongyi-MAI/MAI-UI** — Qwen-UI-Agent: Towards Next-Generation Real-World Centric Foundation GUI Agent

- Repository: https://github.com/Tongyi-MAI/MAI-UI
- Website: https://tongyi-mai.github.io/Qwen-UI-Agent/
- Stars: 2,382 · Forks: 217
- Language: Jupyter Notebook
- License: not declared
- Published: 2026-10-07 · Updated: 2026-10-07 · Language: en
- Canonical page: https://hysenlabs.com/projects/tongyi-mai-mai-ui

## The repository name and the thing it describes are different

The repository is Tongyi-MAI/MAI-UI. The GitHub description reads Qwen-UI-Agent: Towards Next-Generation Real-World Centric Foundation GUI Agent. The README's opening line introduces Qwen-UI-Agent, described as a real-world-centric foundation GUI agent that unifies mobile, computer, browser and DeepSearch scenarios in a single model.

So there are two names for what is being presented here, and the repository structure keeps both. The Projects section lists Qwen-UI-Agent as the continuation work of MAI-UI, and MAI-UI 1.0 as the original repository content, each with its own project website. The tree at the root has exactly two directories, `MAI-UI/` and `Qwen-UI-Agent/`, plus `README.md`.

The README also carries an image labelled as the evolution from MAI-UI 1.0 to Qwen-UI-Agent, which is the clearest statement that this is a successor rather than a rename. For anyone trying to work with this, the practical question is which directory holds what, and the answer is that the older work sits under its own name while the newer work sits in its own.

One more metadata detail is worth carrying. GitHub reports the language as Jupyter Notebook, which tells you the substantive content is notebooks and documents rather than application code. There is no LICENSE file in the tree and no license is reported by GitHub, while the README's table of contents includes a License section. If you intend to build on this work, resolving the licensing is something you need to do directly rather than infer from the repository.

## Reading the score list carefully, because the two lists differ

The contribution list gives eight benchmark results for Qwen-UI-Agent: 82.1% on MobileWorld, 92.2% on MobileWorld-Real, 97.5% on AndroidDaily, 79.5% on OSWorld-Verified, 40.0% partial score on OSWorld-v2, 73.6% on WebArena, and 81.5% on ScreenSpot-Pro.

The dated news entry announcing Qwen-UI-Agent on 2026-07-30 gives a mostly overlapping but differently composed set: the same 82.1% MobileWorld, 92.2% MobileWorld-Real, 97.5% AndroidDaily, 79.5% OSWorld-Verified, 73.6% WebArena and 81.5% ScreenSpot-Pro figures, plus 75.0% on BrowseComp-ZH, and no OSWorld-v2 number at all.

Both lists are in the same README. Neither is marked as approximate or superseded. The practical reading is that the contribution list is the current one and includes OSWorld-v2, while the news entry predates or omits that result and instead reports BrowseComp-ZH. If you are citing a number, cite the contribution list and name the benchmark, because a bare figure like 40.0% means something very different on OSWorld-v2 partial score than on a benchmark scored outright.

The news entry goes further and claims the results are competitive with or surpassing frontier models including Claude Opus 4.8, Gemini 3.1 Pro and GPT-5.6 Sol. That comparison is a project claim rather than something the repository substantiates, since no per-model score table is included. The same applies to the earlier dated records: MAI-UI-235B taking first place on the AndroidWorld leaderboard for pure-vision end-to-end models with 76.7%, and MAI-UI at 32B, 8B and 2B ranking first in all size categories on ScreenSpot-Pro with 67.9%, 65.7% and 57.4%, achieved, the README says, without any zoom-in tricks.

## A hundred physical phones is the expensive part

The contribution most worth reading twice is the first one, because it describes infrastructure rather than a technique. The project builds a live mobile environment of 100 or more physical smartphones covering more than 150 apps, used for task construction, trajectory collection, training and evaluation.

Alongside it is MobileWorld-Real, described as a self-built real-device benchmark with more than 400 tasks across more than 100 apps, and the stated purpose is closing the sim-to-real gap.

This is the axis on which GUI agent research is most often untrustworthy, because a model that succeeds in an emulator frequently fails on a real device with different rendering, different latency and different system dialogs. The gap between a MobileWorld score of 82.1% and a MobileWorld-Real score of 92.2% is odd on its face, since the real-device number is higher, and the README does not explain the relationship between the two benchmarks. What it does say is that the real-device set closes the sim-to-real gap, which implies the simulator-based set is the weaker proxy.

The long-horizon training claim is the other infrastructure-heavy part. Online reinforcement learning is run over trajectories exceeding 100 steps, with roughly 10,000 parallel environments rolling out concurrently to accelerate rollout generation. Sustaining ten thousand concurrent device environments is a different kind of engineering problem from training a model, and it is the part of this repository a reader cannot verify from the outside.

## GUI plus Bash in one action space, with roughly forty percent batching

The action space is described as hybrid, combining GUI operations with direct execution of Bash commands, and additionally emitting multiple actions in a single decision. The README states the empirical finding plainly: in computer-use tasks, CLI commands and GUI clicks emerge as the two dominant action types, and roughly 40% of action outputs are batched.

That finding is the interesting argument here. The assumption behind most GUI agents is that the interface is the only way to act on a computer, which means every task is serialised into clicks and keystrokes, and every step carries model latency and error probability. If a large share of steps turn out to be command-line work anyway, then modelling both in one action space removes a great deal of unnecessary clicking.

Batching is a separate axis with a separate benefit. Emitting several actions per model turn reduces the number of round trips, which matters most on long trajectories where per-step latency dominates. Roughly 40% of outputs being batched means the model is choosing to act on several targets at once about two times in five.

The remaining contributions are about the surrounding system rather than the model. An AutoResearch-style data flywheel has agents construct tasks, environments and verifiers, diagnose failures and plan subsequent iterations, which the README says reduces human effort in capability iteration. A harness layer lets the agent initiate tasks from real-world signals such as a flight-cancellation notification, present decision-ready plans for user confirmation, and execute stateful workflows across mobile and computer.

## A demo you can read closely, and pointers to everything else

The Demo section is where the README stops being a list of claims and starts describing specific tasks, which makes it the most useful part to read for judging what the agent actually does.

The first demo is titled real-device mobile use, with the sub-label recipe research and e-shopping, and the user instruction is concrete: plan to make passion-fruit sour-soup beef tonight, search Douyin for the most-saved photo-and-text post, save it and remember the ingredients, then in the Hema app purchase all the ingredients mentioned in the post except seasonings, select delivery for 18:45 that day, and place the order.

Read that as a task specification and you can see what makes it hard. It spans two applications, it requires extracting information from one to act in the other, it carries a time constraint that is checked rather than merely mentioned, and it includes a negative instruction, everything except the seasonings. The exclusion is the kind of detail that separates a benchmark task from a demo.

The demos themselves are not hosted in the repository. The README directs you to the Qwen-UI-Agent project website for all interactive examples and embeds preview images below. Everything else is also a pointer: a technical report PDF on the project site, an arXiv paper, the MAI-UI project website, and the MobileWorld benchmark site.

Two of those links are worth singling out for what they signal. The technical report is listed as Qwen-UI-Agent-Technical-Report.pdf, and a separate arXiv entry points at 2512.22047 for the original MAI-UI report. MAI-UI-8B and MAI-UI-2B weights were released on Hugging Face at the end of December 2025, so at least the earlier generation of this work is available to download, while the repository tree shows no code for either project.

## Conclusion

This is a research announcement repository, not a tool you install, and the distinction is visible in its shape: three top-level directories, a README carrying most of the content, and no library, package manifest or build entry point at the root. What it does offer is unusually concrete for the genre, because the benchmark is built on real phones rather than a simulator, which is the axis where GUI agent results usually diverge most from practice. The scores need reading with care, since the headline list and the dated news entries do not report identical sets. Start with the technical report on arXiv for method and the project website for interactive demos, and treat the repository as the index that points to both.

## FAQ

### What is Qwen-UI-Agent?

It is a foundation GUI agent that handles mobile, computer, browser and DeepSearch scenarios in one model. The README describes a hybrid action space combining GUI operations with direct Bash execution, batched actions in about 40 percent of outputs, and online reinforcement learning over trajectories longer than 100 steps.

### Is MAI-UI the same project as Qwen-UI-Agent?

Qwen-UI-Agent is the continuation of MAI-UI, not a rename of it. The repository keeps both under separate top-level directories, MAI-UI/ and Qwen-UI-Agent/, and the README describes the relationship as an evolution from MAI-UI 1.0 to Qwen-UI-Agent, with each having its own project website.

### Can I run or install anything from the MAI-UI repository?

Not from the repository itself. The tree contains MAI-UI/, Qwen-UI-Agent/ and README.md, with no package manifest, build entry point or library at the root. The model weights that were released are MAI-UI-8B and MAI-UI-2B on Hugging Face, not a runnable application.

### What is MobileWorld-Real?

It is a real-device benchmark the project built itself, with more than 400 tasks across more than 100 apps, paired with a live environment of over 100 physical smartphones covering 150 or more apps. Its stated purpose is closing the sim-to-real gap that GUI agent benchmarks usually have.

## Sources

- [Issues](https://github.com/Tongyi-MAI/MAI-UI/issues)
- [Project website](https://tongyi-mai.github.io/Qwen-UI-Agent/)
- [README](https://github.com/Tongyi-MAI/MAI-UI/blob/main/README.md)
- [Tongyi-MAI/MAI-UI on GitHub](https://github.com/Tongyi-MAI/MAI-UI)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/tongyi-mai-mai-ui
