# Mobile-Agent: Alibaba's GUI Agent Family, From Mobile-Agent-v1 to GUI-Owl-1.5

> Mobile-Agent is Tongyi Lab's family of GUI agents for phones, desktops and browsers, distributed as a monorepo of versioned directories plus the GUI-Owl model weights. The repo is a research release first and a product second, and the README says so by pointing you at hosted demos before it points you at code.

**X-PLUG/MobileAgent** —  Mobile-Agent: The Powerful GUI Agent Family

- Repository: https://github.com/X-PLUG/MobileAgent
- Stars: 9,265 · Forks: 925
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/x-plug-mobileagent

## What Mobile-Agent Actually Is, and Who It Is Built For

Mobile-Agent is not one program. It is a repository that holds several generations of a GUI agent research line from Tongyi Lab at Alibaba Group, each in its own top-level directory: Mobile-Agent-v1, Mobile-Agent-v2, Mobile-Agent-v3, Mobile-Agent-v3.5, Mobile-Agent-E, PC-Agent, UI-S1 and GUI-Critic-R1. The README describes the project as "Mobile-Agent: The Powerful GUI Agent Family by Tongyi Lab, Alibaba Group" and the underlying models as native multi-platform GUI agent foundation models covering desktop, mobile and browser automation.

The intended reader is a researcher or an engineer who needs an agent that reads a screen and emits taps, clicks and text rather than one that calls an API. The README's first call to action is not an install command. It is a link to a ModelScope online demo and a Bailian online demo for Mobile-Agent-v3.5, followed by a note that a Mobile-Agent-v3.5 API is available on Bailian. That ordering tells you who the maintainers expect to show up first: people who want to see the behaviour before they commit a GPU.

If you are looking for a pip-installable library that automates a phone in ten lines, this repository will frustrate you. If you are looking for the code and weights behind a published agent family, with versioned snapshots you can pin, it is unusually complete. The MIT licence covers the repository; the model weights are distributed separately through HuggingFace and ModelScope collections, and their terms are not stated in the README.

## How a GUI Agent Works Here: Screenshots In, Actions Out

The mechanism across the family is the same shape. A multimodal large language model receives a screenshot of the current interface plus a task instruction, and produces an action, typically a tap or click at a coordinate, a text entry, or a scroll. The agent then observes the resulting screen and repeats. What differs between versions is how much scaffolding sits around that loop.

The v3.5 line is built on GUI-Owl-1.5, described in the release notes as a family of native multi-platform GUI agent foundation models built on Qwen3-VL, released in sizes 2B, 4B, 8B, 32B and 235B, in both Instruct and Thinking variants. The README states these support desktop, mobile and browser automation and reports state-of-the-art results on more than 20 GUI benchmarks, with emphasis on end-to-end tasks, grounding, tool and MCP calling, and long-horizon memory. Those are the maintainers' claims in the repository, not measurements performed here.

The repository layout reflects an architectural split rather than a single pipeline. Mobile-Agent-E and GUI-Critic-R1 sit beside the versioned agents, and UI-S1 is described as work on semi-online reinforcement learning for GUI automation, with its own paper, code, dataset and a 7B model. ToolCUA, announced in the news list, is described as an end-to-end Computer Use Agent for GUI and tool path orchestration, trained in two stages: trajectory-aware tool synthesis, then online agentic reinforcement learning. That is a different bet from pure screenshot-to-coordinate agents. It assumes some actions are better served by calling a tool than by driving the interface.

Long-horizon memory is the part worth scrutinising. The README lists it as a capability of GUI-Owl-1.5 but does not describe the memory representation, its size limits, or what happens when a task exceeds them. For multi-step mobile workflows, that is the detail that decides whether the agent finishes or loops.

## Installing Mobile-Agent: Where the README Sends You First

There is no top-level installation command in the README. The repository is a monorepo, and each version directory carries its own README with its own setup. The README does give two paths that avoid local deployment entirely, and for a first evaluation those are the honest starting points.

The README states you can try Mobile-Agent-v3.5 through the ModelScope online demo at modelscope.cn/studios/MobileAgentTest/computer_use, or through the Bailian online demo. It also states that a Mobile-Agent-v3.5 API is provided on Bailian, with documentation at the Bailian console under the model gui-plus-2026-02-26. If you only want to know whether the agent can drive your workflow, that page is where to look.

For the model weights themselves, the README points to the GUI-Owl-1.5 collections on HuggingFace and ModelScope, and separately to GUI-Owl-32B and GUI-Owl-7B model pages. The 7B and 32B pages are the concrete artefacts you would pull for a self-hosted run:

```bash
# Model pages named in the README
git clone https://huggingface.co/mPLUG/GUI-Owl-7B
git clone https://huggingface.co/mPLUG/GUI-Owl-32B
```

Cloning those repositories downloads the weights and the model card files. What you should see afterwards is a directory containing the model configuration and weight shards, not a runnable agent. The agent code lives in the Mobile-Agent-v3.5 directory of this repository, and its own README is the place to look for the runtime instructions, since the top-level README does not repeat them.

For the mobile path, the README's news entry for 2025.9.16 says the code of GUI-Owl and Mobile-Agent-v3 on OSWorld, AndroidWorld and real-world mobile scenarios has been open-sourced, and links to a section titled "deploy Mobile-Agent-v3 on your mobile device" inside the Mobile-Agent-v3 directory. That is the pointer to follow for device deployment. There is also a Wuying Cloud Phone route announced on 2026.3.31, described as a cloud-based Android environment for running Mobile-Agent-v3.5, with a link to Alibaba Cloud documentation. That option removes the need to own an Android device, at the cost of running inside Alibaba's infrastructure.

## Limits You Should Know Before Building On It

The strongest limitation is the one the repository structure advertises: there is no stable, unified interface. Each directory is a snapshot of a research iteration with its own dependencies and its own README. Upgrading from Mobile-Agent-v2 to v3 is not a version bump, it is a migration between projects that happen to share a repository. The top-level README does not document rollback, deprecation policy, or which versions continue to receive fixes.

Second, the top-level README never states hardware requirements for self-hosting. It lists model sizes up to 235B but does not say what runs where, so the decision about which GUI-Owl-1.5 variant is feasible is left to you. That is a real planning gap, not a small one.

Third, the README's own framing is cloud-first. The recommended experiences are hosted demos, a Bailian API, and Wuying Cloud Phone. A team that must keep screen data on its own hardware has to work from the per-version READMEs, and the top-level document gives no privacy or data-handling statement for the self-hosted path. Screenshots of a phone contain whatever is on that phone.

Finally, evaluation claims are reported as benchmark results, not as field reliability. The README says GUI-Owl-1.5 achieves state-of-the-art results on more than 20 GUI benchmarks. Benchmarks do not tell you how the agent behaves on an app that renders its controls in a canvas, or on a screen with a modal that appeared mid-task. Treat the numbers as the authors' evidence for model quality, and test your own apps.

## Mobile-Agent Versus Appium and Playwright

The natural comparison for anyone doing automation is Appium for mobile or Playwright for browsers. The difference in approach is fundamental. Appium and Playwright drive applications through their accessibility trees, DOM selectors, or platform automation APIs. A test written against them refers to a button by an identifier or a selector, and it breaks when that identifier changes.

Mobile-Agent drives the interface the way a person does: it looks at pixels and issues coordinates. That means it can operate an app whose internals you cannot address, including third-party apps and screens rendered by a game engine or a canvas. It also means its correctness depends on visual grounding, and a mispredicted coordinate is a wrong action with no selector to blame. The README's framing around grounding and end-to-end tasks reflects exactly this trade-off.

There is a middle path in the repository itself. ToolCUA is described as orchestrating GUI actions and tool calls, with a training pipeline designed to learn when to use each. If your workflow has a mix of well-defined API calls and stubborn interfaces, that is the branch of the family aimed at you, though the README only announces it and links to a separate GitHub repository.

For deterministic regression tests against your own application, Appium or Playwright remain the better tool. Mobile-Agent is for the interfaces you do not control.

## Maintenance, Licence and the Cost of Keeping Up

The repository is not archived, and the last push was on 2026-07-07. The news list in the README shows a steady cadence of releases across late 2025 and the first half of 2026: GUI-Owl and Mobile-Agent-v3 in August 2025, UI-S1 in September, GUI-Owl-1.5 in February 2026, its online inference availability in March, and ToolCUA in May. That cadence is the upgrade cost. New model generations arrive roughly every few months, and each generation is a new directory rather than a new tag.

The licence is MIT, which covers the code in this repository. The README does not state the licence of the GUI-Owl model weights; they live in separate HuggingFace and ModelScope repositories whose own terms govern them. If you plan to redistribute a product built on those weights, check the model repositories directly rather than assuming MIT carries over. Nothing here is legal advice.

Operationally, budget for two things the README does not quantify: the compute to host a GUI-Owl-1.5 model of the size you choose, and the work of pinning a version directory and reading its README before every bump. Because the versions are directories, a git submodule or a sparse checkout of one directory is a more honest dependency than tracking main.

## Conclusion

Adopt Mobile-Agent if you want to read or reproduce GUI agent research, or if you are building on Alibaba Cloud and can use Bailian, ModelScope or Wuying Cloud Phone as the runtime. Do not adopt it as a drop-in library: there is no single package, no PyPI name in the README, and the entry points are per-version directories. Before committing, verify three things yourself: which subdirectory matches your target platform, whether the GUI-Owl-1.5 model size you can host matches the 2B/4B/8B/32B/235B list, and whether your device or emulator exposes the ADB or desktop control channel that the version you picked expects.

## FAQ

### What is Mobile-Agent?

It is a family of GUI agents from Tongyi Lab at Alibaba Group, released as a repository containing several versioned agent directories plus the GUI-Owl model line. The agents read a screen and produce interface actions such as taps, clicks, text entry and scrolling.

### Can you give me an example of a mobile agent?

Mobile-Agent-v3.5 is the current example in this repository, built on the GUI-Owl-1.5 models. The README points to a ModelScope demo at modelscope.cn/studios/MobileAgentTest/computer_use and a Bailian demo where you can enter an instruction and watch the agent operate an interface.

### How do you install Mobile-Agent?

There is no top-level install command; each version directory has its own README. The README's fastest routes are the hosted ModelScope and Bailian demos, or the Mobile-Agent-v3.5 API on Bailian, while the model weights are pulled from the GUI-Owl-1.5 collections on HuggingFace and ModelScope.

## Sources

- [Issues](https://github.com/X-PLUG/MobileAgent/issues)
- [License: MIT](https://github.com/X-PLUG/MobileAgent/blob/main/LICENSE)
- [README](https://github.com/X-PLUG/MobileAgent/blob/main/README.md)
- [X-PLUG/MobileAgent on GitHub](https://github.com/X-PLUG/MobileAgent)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/x-plug-mobileagent
