Model or dataset
Purewhiter/mobilegym avatar
Purewhiter/mobilegym

MobileGym and the 102 GB of RAM hiding behind one six minute evaluation

[EMNLP 2026] MobileGym: A Verifiable and Highly Parallel Simulation Platform for Mobile GUI Agent Research · 浏览器里运行的安卓模拟器 · Browser-hosted Android Simulator · Verifiable Evaluation · Scalable Online RL Training

803 stars137 forksPythonApache-2.0

At a glance

What is it?
A browser-hosted Android simulation environment with 28 simulated apps, 416 task templates and programmatic judges instead of VLM judging. Its benchmark level 1 is saturated at twenty tasks while level 4 decides the ranking, and every map is a bundled snapshot with an optional browser-visible API key.
Who is it for?
Read the resource figures as a cluster requirement rather than a laptop one. Four hundred megabytes per instance times 256 instances is a hundred gigabytes before anything else runs, and the claim of a six minute full evaluation is a claim about that machine, not about your workstation.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 14 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.

Editorial analysis

400 MB times 256 is a hundred gigabytes before anything else runs

The headline figure is that a full 256-task evaluation finishes in about six minutes, with 256 parallel instances on one server using under 10% CPU. The per-instance numbers are given just as plainly: roughly 400 MB of RAM and roughly 50 MB of disk each, and about 3 seconds of cold start per instance.

Multiply the memory and the shape of the requirement changes. Four hundred megabytes times 256 instances is a little over 102 GB of resident memory, plus 12.8 GB of disk, before the host, the benchmark harness or the agent under test. The CPU claim is about the same machine, and a simulation built on deterministic state checks is exactly the kind of workload where idle CPU and resident memory trade against each other, so a hundred gigabytes of RAM is not an odd way to buy the six minutes.

The last push to the default branch is dated 21 September 2026, so the figures are recent. They are also, as far as this file goes, unattributed to any particular machine. No CPU model, no core count, no operating system appears alongside them.

Level 1 is saturated at twenty tasks and level 4 decides the ranking

The benchmark is called MobileGym-Bench with 256 test tasks, split across four levels that carry their own sample sizes: L1 at 20, L2 at 73, L3 at 83 and L4 at 80. Those add up to the 256, so the split is a partition rather than a sample of one.

The two visible leaderboard rows show where the difficulty sits. Gemini 3.1 Pro scores 97.5 on L1 and Doubao-Seed-2.0-Pro scores 100.0, and Qwen3.6-Plus also 100.0. By L4 the same two models are at 21.9 and 6.2. So the easiest fifth of the benchmark, twenty tasks, is effectively finished for every frontier model listed, and the eighty hardest tasks separate a 58.8 percent overall from a 52.0 percent by a factor of three and a half.

A ranking built that way is really an L4 ranking with a large block of saturated tasks diluting the average. Two other columns, PR, FC and USE, appear in the table with values such as 72.1, 34.0 and 5.5, and neither their meaning nor the definition of the four levels is given in the part of the file that carries the table.

Two license files, and one of them covers the data

The header carries two licence badges pointing at two different files, `LICENSE` and `LICENSE-DATA`. The tree confirms both sit at the root, next to `NOTICE` and a `DISCLAIMER.md`.

Splitting a repository's licence from its data is a common and defensible choice, since a benchmark's task content and assets often carry terms that code should not inherit. What this file does not do is say which is which. Both badges look identical in the header, the description field does not mention either, and nothing in the visible text explains whether the simulated apps, the task templates, the map snapshots and the captured resources are under the data licence or the code one.

For a research platform whose selling point is a 28-app catalogue and 416 task templates, that is the licence question worth answering first, and it takes opening two files to answer rather than one.

Every VITE key is visible in browser JavaScript, so every map is a snapshot

The environment template opens with a warning that `VITE_*` values used by frontend code are visible in browser JavaScript, and that browser-side keys should be protected with referrer or domain restrictions. Every key in the file is prefixed that way, and there are three of them for third-party services.

`VITE_GOOGLE_MAPS_API_KEY` is described as recommended for the best map rendering and live fallback but optional for the canonical benchmark split. Without it, the Map app uses bundled places and routes snapshots first and can load captured Maps JavaScript and resources from a local Service Worker cache. Benchmark tasks stay usable, and what you lose is uncached places, details, tiles and online fallbacks. `VITE_GOOGLE_MAP_ID` is explicitly not a secret, and an empty value makes the app use its demo map ID instead. `VITE_AMAP_API_KEY` is for reverse geocoding, and the template is careful about the failure: with no key, direct reverse geocode calls fail, the Weather app catches that and substitutes an offline name if it has one, and otherwise the location name comes back generic or empty.

That is a benchmark designed to be reproducible without a network, with three degradation paths written down in advance. It is also three third-party browser-visible keys on the optional path.

SQLite compiled to WebAssembly is a runtime dependency, not a build tool

The mechanism behind the no-install claim is one line in the dependency list: `@sqlite.org/sqlite-wasm` at ^3.53.0-build1, and it sits in `dependencies` rather than in `devDependencies`. A WebAssembly build of SQLite is the storage engine of the running application, not something a bundler touches once and throws away.

The rest of the root manifest describes a Vite and React front end on React 19, with `react-router-dom` 7.10.1 and `zustand` 5.0.11 for state, `immer` for updates, `leaflet` and `@googlemaps/js-api-loader` for maps, `lucide-react` for icons, `qrcode` for codes, and `@tanstack/react-virtual` for long lists. Dev-side, `puppeteer` at ^24.34.0 and `vitest` at ^4.0.18 sit next to `typescript` pinned with a tilde at ~5.8.2, `vite` at ^6.2.0, `eslint` at ^10.0.2 and `typescript-eslint` at ^8.56.1.

Two things stand out in that list. The storage engine is the one component with a build suffix in its version, so it moves on a different cadence from everything around it. And the lint stack is on a new major while the TypeScript plugin is on 8.x, which is the pairing most likely to need attention first.

The lint script covers two directories out of the ones at the root

One script defines the quality gate: `eslint os/ apps/ && node scripts/lint_store_getters.mjs`. Two directories, plus a custom rule that is not part of ESLint at all.

The root listing has more code roots than that. Alongside `os/` and `apps/` there is `system/`, `bench_env/`, `web/`, `tests/`, `scripts/`, `mobilegym-rl/`, `assets/`, `public/`, `docs/` and `.nginx/`, and none of them is named in the lint command. The CSS has its own pair of scripts that run the Tailwind CLI against `web/tailwind.input.css` to produce a minified `web/tailwind.css`, and the tests run through `vitest run` with a watch variant.

The store-getter script is the more interesting half. A hand-written check on store getters is a convention that ESLint will not enforce for you, and its presence says the state layer has a rule that is worth a script. It is also a rule that only applies to the two directories the script covers, so the convention is local to the simulator core rather than repository-wide.

The table of contents promises five quick start sections this file never reaches

The contents list in the file names five numbered steps: install, boot the simulator, talk to an agent in plain language, run the benchmark, and train with RL. It also promises sections on how it works, the apps catalogue, architecture at a glance, extending MobileGym, and a citation block. The table of contents for extending and citing is as detailed as the quick start's.

This version of the file reaches the leaderboard table and stops partway through the third row, so none of those five steps has a command in it, and the architecture, the app catalogue and the 28-app manifest contract are all named in the contents without appearing. The two releases are the other half of the picture: `v0.1.0` on 26 June 2026 and `data-v0.1.0` two and a half minutes later, with data versioned as its own tag, while the news section dates the release announcement to 27 June.

What is complete is the design argument and the measurement setup, which is enough to judge the idea and not enough to run it. For the environment itself, the template file is the only setup detail written down, and it points at three API keys and nothing else.

95.1 percent of a 42.8 point gain, on one device model

The transfer claim is specific enough to check. A GRPO run on Qwen3-VL-4B gains 42.8 points in simulation, and 95.1 percent of that gain is retained on a real device, which is 40.7 points. The arithmetic holds. The real device is a Redmi Note 12 Turbo, and the framing is deliberate: behavioural fidelity, not pixel fidelity.

The supporting case study is narrower and older, dated April 2026, and says so. It reports a 40.7 point real-device gain on a signal bucket tasks subset after 10 GRPO steps on one node. So the headline retention figure and the case study share a number, and the case study is a subset, a fixed step count and a single node.

The surrounding claims are consistent with that. Group-RL methods such as GRPO need state that can be reset and cloned, which is the second of the three walls the project sets up: real app state hides in encrypted databases and server backends, so it can be neither reset nor cloned. The 10.2 percent VLM misjudgment figure in the same table is offered as the cost of the alternative, verification by asking a model whether the screen looks right.

Editorial conclusion

Read the resource figures as a cluster requirement rather than a laptop one. Four hundred megabytes per instance times 256 instances is a hundred gigabytes before anything else runs, and the claim of a six minute full evaluation is a claim about that machine, not about your workstation. The evaluation design is the strongest part of the project: programmatic check functions per task, a state snapshot that judges read directly, and full-environment diffing for the side effects a real device cannot show you. The gaps to check before you rely on it are narrower and worth naming. The level definitions for L1 through L4 and the columns labelled PR, FC and USE are not explained in the visible part of the file. Data is licensed separately from the code under LICENSE-DATA. And the quick start exists in the table of contents but its commands are not in this version, so the environment template is the only setup detail written down here.

Frequently asked questions

What is MobileGym?

A browser-hosted Android simulation environment for mobile GUI agent research, with fully programmable state kept as a structured JSON snapshot. It ships 28 simulated apps and 416 task templates, and every task has a programmatic check function so verification does not depend on a VLM judge. Its paper was accepted to the EMNLP 2026 main conference.

How is MobileGym different from a real device or an emulator?

The stated argument is that adb and accessibility trees expose the interface but not balances, orders or chat history, so verification falls back on stochastic VLM judges, with 10.2 percent misjudgment measured. MobileGym keeps the environment as structured JSON, so judges read state directly, and state can be reset, injected, snapshotted, cloned and rolled back across hundreds of parallel instances.

How much hardware does one MobileGym evaluation need?

About 400 MB of RAM and 50 MB of disk per instance, with roughly 3 seconds of cold start each. The file claims 256 parallel instances on a single server using under 10 percent CPU, and a full 256-task evaluation finishing in about 6 minutes. No CPU model or host specification is given alongside those figures.

Does MobileGym need Google Maps or weather API keys to run?

No. The environment template calls VITE_GOOGLE_MAPS_API_KEY recommended for live fallback but optional for the canonical benchmark split, because bundled places and routes snapshots plus a local Service Worker cache cover what tasks need. VITE_AMAP_API_KEY is optional too, with an empty value making reverse geocoding calls fail and the Weather app substitute an offline name.

Official sources

  1. License: Apache-2.0
  2. Project website
  3. Purewhiter/mobilegym on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/purewhiter-mobilegym.svg)](https://hysenlabs.com/projects/purewhiter-mobilegym)