# GELab-Zero: a locally deployed GUI agent stack for Android phones

> GELab-Zero pairs a 4B GUI agent model with the engineering plumbing that usually blocks mobile agents: ADB device management, task distribution and trajectory recording. The trade-off is a narrow target: it is built for Android, not desktop.

**stepfun-ai/gelab-zero** — STEP-GUI: The top GUI agent solution in the galaxy.  Developed by the StepFun-GELab team and powered by StepFun’s cutting-edge research capabilities.

- Repository: https://github.com/stepfun-ai/gelab-zero
- Website: https://opengelab.github.io/
- Stars: 2,276 · Forks: 199
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/stepfun-ai-gelab-zero

## The engineering tax GELab-Zero is trying to remove

Most published GUI agent work stops at the model. The README is explicit that the hard part is elsewhere: multi-device ADB connections, dependency installation, permission configuration, inference service deployment, and task recording and replay. Anyone who has tried to run a phone-use agent knows this list. A research demo works on one handset with one cable. The moment you want five phones running the same task set, you are writing device discovery, retry logic and log capture yourself.

GELab-Zero's stated audience is two groups. Agent developers who want to test new strategies and interaction approaches without rebuilding the harness each time, and enterprise users who want to integrate MCP capabilities into a product. The repository layout backs this up: copilot_agent_client, copilot_agent_server, copilot_front_end and copilot_tools sit alongside model_config.yaml and mcp_server_config.yaml, which reads as a client/server split with a configuration layer rather than a single script. The README describes the project as a fully open-source GUI agent with both model and infrastructure, deployable entirely locally. That pairing is the actual pitch: weights plus plumbing.

## Model, server, client: how the pieces are arranged

The architecture visible in the repository is a set of cooperating processes. A model layer, configured through model_config.yaml, serves the GELab-Zero-4B-preview weights. A server component under copilot_agent_server handles orchestration. A client under copilot_agent_client drives the device. Tools live in copilot_tools and tools/. The MCP path is separate and explicit: mcp_server_config.yaml configures an MCP server, and the README's news entry for 2025-12-12 shows how to start it for multi-device management and task distribution, then import those tools into Chatbox.

The README names three agent modes: ReAct loops, multi-agent collaboration, and scheduled tasks. It also states that tasks can be distributed to multiple phones while interaction trajectories are recorded, which is what makes the setup reproducible rather than a one-off demo. Device control goes through uiautomator2, listed in requirements.txt, which is the standard Python route into Android's UI automation layer over ADB. There is also a yadb entry at the repository root, which the README does not explain.

The 4B size is the design constraint that shapes everything else. The README claims it runs on consumer-grade hardware and balances low latency with privacy. It does not publish latency figures or hardware requirements, so treat that as a target rather than a measured property.

## Installing GELab-Zero and running a first task

The README points to a project page at opengelab.github.io and a Chinese quick-start document, 极简运行指南_CN.md, at the repository root. There is no English install section in the README text itself, so the dependency list in requirements.txt is the concrete starting point. It pulls in the model-serving and device layers together: openai==2.28.0 for the inference client, uiautomator2 for the phone, fastmcp for the MCP server, and streamlit for the front end.

```bash
pip install -r requirements.txt
```

After that, the weights come from Hugging Face or ModelScope. The README links stepfun-ai/GELab-Zero-4B-preview on both, and separately thanks third parties for GGUF and EXL3 quantizations. Those quantized builds are community artifacts, not part of the repository, so the format you pick determines which runtime you need.

The MCP server is the documented way to get multi-device management and task distribution. The README gives this exact command:

```bash
# enable mcp server
python mcp_server/detailed_gelab_mcp_server.py
```

The README's next step is to import the MCP tools into Chatbox, with a screenshot showing the result. If you prefer to drive things from Python, the examples directory contains run_single_task.py, run_single_task_state_compress.py, run_task_via_mcp.py and run_test_api.py. The README does not document the arguments for any of them, so read the files before running them. Expect any of them to need a connected device and a reachable inference endpoint; neither the README nor requirements.txt states the default port for the model service.

## Where GELab-Zero stops being the right tool

Android only. Every device-control dependency in requirements.txt is Android-flavored, and the demonstrations are phone tasks: finding a sci-fi movie, picking a weekend destination for kids, claiming meal vouchers on an enterprise welfare platform. If your target is a desktop application, a browser, or an iOS device, this stack gives you nothing.

The 4B model is the second boundary. A small model that runs on consumer hardware is a deliberate trade: it keeps inference local and cheap, and it will lose to a larger hosted model on long, ambiguous task chains. The README does not publish success rates or compare the 4B model against larger alternatives, so you cannot size that gap from the documentation.

Operational fragility is the third. GUI agents act on pixels and tap coordinates, so a redesigned app screen or a new permission dialog can break a task that worked yesterday. The trajectory recording helps you diagnose that, but it does not prevent it. Finally, maintenance: the last push was on 2026-05-11, and the README's news list opens with a placeholder item marked Coming Soon. There are no retrieved releases, so plan on tracking the main branch rather than pinning a version.

## How this differs from hosted phone-agent APIs

The obvious alternative is a hosted mobile automation service: you send screenshots or accessibility trees to a vendor endpoint, the vendor runs a larger model, and you get back actions. The difference is not just where the model runs. With a hosted API you inherit the vendor's device fleet and their model upgrades, and you give up custody of the screen content. GELab-Zero inverts that. The README's privacy claim, that deployment is entirely local with no cloud dependencies, is only meaningful if you actually run the 4B weights on your own machine, which means you own the GPU or CPU budget and the model quality ceiling.

A second comparison is the MCP route. GELab-Zero exposes its device management and task distribution as MCP tools, which means an existing MCP client such as Chatbox can drive phones without you writing an agent loop. That is a different integration shape from a Python SDK: less control over the loop, far less code to write. The README shows the Chatbox path but does not document the tool schemas, so the practical difference is that you are trusting the client's orchestration instead of your own.

## Licence and the cost of staying current

The repository is MIT licensed, with an additional Notice.txt at the root. MIT is permissive: it allows commercial use and modification with attribution and no warranty. The Notice.txt file is worth reading before you ship, since it may carry terms that sit alongside the licence, and nothing in the README explains what it contains. This is not legal advice; if the deployment matters commercially, have someone read both files.

Upgrade cost is the more practical question. With no retrieved releases and a last push on 2026-05-11, there is no versioned artifact to pin. The model itself lives outside the repository, on Hugging Face and ModelScope, so model updates and code updates can drift apart. The community GGUF and EXL3 conversions add a third track: each quantization targets a different runtime, and the README credits them to third-party authors rather than maintaining them. Keeping a working setup means tracking three things that move independently.

## Conclusion

GELab-Zero fits teams that already have Android devices on ADB and want the agent loop, device management and trajectory logging handled locally, without sending screen data to a hosted service. It is the wrong pick if you need desktop GUI automation, a managed cloud endpoint, or a model larger than 4B. Before adopting it, verify three things in your own checkout: that the 4B weights download from Hugging Face or ModelScope, that your phone enumerates over ADB, and that the entry points under examples/ run against your device. The last push to the repository was on 2026-05-11, so treat the code as a snapshot rather than a stream of fixes.

## FAQ

### What is GELab-Zero?

It is a GUI agent system for Android from StepFun's GELab team, combining a 4B model with the inference infrastructure around it. The README describes it as the first fully open-source GUI Agent with both model and infrastructure, deployable entirely locally.

### How do I install GELab-Zero?

Install the Python dependencies listed in requirements.txt, then obtain the GELab-Zero-4B-preview weights from Hugging Face or ModelScope. The README also points to a Chinese quick-start guide, 极简运行指南_CN.md, at the repository root.

### Does GELab-Zero need a cloud service?

The README states there are no cloud dependencies and that deployment is entirely local, which is how it frames its privacy control. In practice that means you host the 4B model yourself, so the hardware budget moves to your side.

## Sources

- [Issues](https://github.com/stepfun-ai/gelab-zero/issues)
- [License: MIT](https://github.com/stepfun-ai/gelab-zero/blob/main/LICENSE)
- [Project website](https://opengelab.github.io/)
- [README](https://github.com/stepfun-ai/gelab-zero/blob/main/README.md)
- [stepfun-ai/gelab-zero on GitHub](https://github.com/stepfun-ai/gelab-zero)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/stepfun-ai-gelab-zero
