# Open-AutoGLM: an ADB-driven phone agent you run from a Python CLI

> Open-AutoGLM pairs a vision-language model with ADB or HDC to operate Android and HarmonyOS phones from natural-language instructions. It is an alpha-stage research framework, not a consumer app.

**zai-org/Open-AutoGLM** — An Open Phone Agent Model & Framework. Unlocking the AI Phone for Everyone

- Repository: https://github.com/zai-org/Open-AutoGLM
- Website: https://autoglm.z.ai/blog
- Stars: 26,312 · Forks: 4,041
- Language: Python
- License: Apache-2.0
- Published: 2026-09-21 · Updated: 2026-09-21 · Language: en
- Canonical page: https://hysenlabs.com/projects/zai-org-open-autoglm

## What Open-AutoGLM does that a scripted UI test cannot

A conventional Android automation script hard-codes resource IDs and tap coordinates. Change the app version and it breaks. Open-AutoGLM takes a different route: it reads the screen as an image, asks a vision-language model what to do next, and executes the resulting action over ADB. The README describes the loop as multimodal screen understanding plus planning plus execution, with the user supplying only a sentence such as opening an app and searching for something. The intended audience is developers who want an agent that survives small UI changes, not QA engineers who need deterministic regression tests. Two design choices matter here. First, the framework keeps a human in the loop for sensitive steps and for login or CAPTCHA screens, which the README lists as built-in behaviour. Second, it supports remote ADB over WiFi or a network connection, so the phone does not have to sit tethered to the machine running the agent.

## The control loop: screenshots in, ADB actions out

The architecture visible in the repository is thin by design. main.py is the entry point and phone_agent/ holds the framework code, with examples/basic_usage.py and examples/demo_thinking.py showing how the pieces are called. The agent talks to the model through the OpenAI-compatible client listed in setup.py (openai>=2.9.0), which means any endpoint that speaks the OpenAI chat format can serve as the brain. On the device side it shells out to ADB for Android and HDC for HarmonyOS NEXT, and ios.py plus docs/ios_setup/ios_setup.md cover the iPhone path through WebDriverAgent. Pillow>=12.0.0 is a dependency because screenshots must be processed as images before they go to the model. That is the whole data flow: capture, encode, send to the model with the current task, receive an action, translate it into an ADB or HDC command, repeat. The README does not document a rollback or undo mechanism, so an action that changes app state stays changed.

## Installing Open-AutoGLM and running a first instruction

The README assumes Python 3.10 or newer. Install the package and its dependencies from the repository root. The second command installs the local package itself, since the project is not published to PyPI under this name.

```bash
pip install -r requirements.txt
pip install -e .
```

Before the agent can do anything, the phone must be reachable. Enable developer mode and USB debugging on the device, connect a data-capable cable, and confirm the device is listed. On Android the README uses adb devices; on HarmonyOS NEXT the equivalent is hdc list targets.

```bash
adb devices
```

The output should show a serial number followed by the word device. If the list is empty, the README suggests the connection failed and that some phones need a restart after enabling developer options. Android devices also need the ADB Keyboard app installed and enabled, either through the keyboard settings or with adb shell ime enable com.android.adbkeyboard/.AdbIME; HarmonyOS devices use their native input method instead. You can skip local model hosting entirely by pointing the CLI at a hosted endpoint. The README gives the BigModel and ModelScope variants, and the model name differs between them.

```bash
python main.py --base-url https://open.bigmodel.cn/api/paas/v4 --model "autoglm-phone" --apikey "your-bigmodel-api-key" "打开美团搜索附近的火锅店"
```

With a valid key, the agent should begin reading the screen and issuing taps and swipes. If you prefer to host the model, the README documents vLLM and SGLang launch commands on port 8000 with --served-model-name autoglm-phone-9b, and notes that the model is structurally close to GLM-4.1V-9B-Thinking.

## Where the framework gets in your way

The setup.py classifier reads Development Status :: 3 - Alpha, and the README repeats that the project is for research and study only. Treat the API as unstable. The model choice is the sharpest constraint: AutoGLM-Phone-9B is tuned for Chinese phone apps, while AutoGLM-Phone-9B-Multilingual adds English, so an app whose interface is in a third language is outside what the README claims. Self-hosting is not lightweight. The vLLM command sets --max-model-len 25480 and --limit-mm-per-prompt with image set to 10, meaning each step can carry up to ten images; the SGLang command uses a context length of 25480 and a max_pixels of 5000000. Those are GPU-shaped numbers. The README also warns that dependency conflicts around transformers can be ignored, which is honest but also a sign that the deployment path is not polished. Finally, the human takeover for login and CAPTCHA is a feature, yet it also means unattended runs will stall at exactly those screens. If your task requires no human presence at all, this is the wrong tool.

## Open-AutoGLM against Midscene.js

The README names Midscene.js as an integration rather than a rival, but the two solve the same problem from opposite ends. Midscene.js is a JavaScript and YAML UI automation SDK driven by vision models, and it has adapted the AutoGLM model so that you can drive iOS and Android through its own flow syntax. The difference is where the orchestration lives. Open-AutoGLM keeps the loop in Python inside phone_agent/, with main.py as the CLI and the OpenAI client as the model boundary; you extend it by writing Python. Midscene.js keeps orchestration in JavaScript or declarative YAML, which suits teams whose test or automation stack is already Node-based. If your team writes Python and wants to inspect or modify the agent loop, Open-AutoGLM is the more direct fit. If you want to describe flows in YAML and reuse an existing JS toolchain, going through Midscene.js to reach the same model is the shorter path.

## Licence, maintenance and the cost of staying current

The repository is licensed Apache-2.0, which permits commercial use and modification provided you keep the licence and notice files; the LICENSE file at the repository root is the authoritative text, and the README's own terms of use in resources/privacy_policy.txt sit alongside it. Note that the README restricts the project to research and study and forbids illegal use, so the licence grant and the project's stated intent are not identical documents and both deserve a read. The repository is not archived, and its last push was on 2026-03-06. Nothing in the repository describes a release process, versioned changelogs or a migration path between versions, so upgrading means tracking main and re-reading the README. The heaviest ongoing cost is the model side: either you pay per call against BigModel or ModelScope, or you maintain a GPU deployment with the pinned vLLM or SGLang parameters, which is where most of your operational time will go.

## Conclusion

Open-AutoGLM fits developers who want to prototype phone automation with a vision-language model and already have ADB or HDC experience, plus a model endpoint or GPU capacity for a 9B model. It does not fit anyone wanting a packaged app: the repository ships no APK, the README points to the community-built ADB Keyboard for text entry, and setup.py still labels the package Alpha. Before committing, verify that your device appears in adb devices or hdc list targets, that you can obtain an API key for autoglm-phone or ZhipuAI/AutoGLM-Phone-9B, and that you accept the terms of use in resources/privacy_policy.txt.

## FAQ

### Does Open-AutoGLM work on iPhone?

Yes, but through a different stack. The README points to docs/ios_setup/ios_setup.md, which covers configuring WebDriverAgent and the iPhone so that AutoGLM can drive it. Android and HarmonyOS use ADB and HDC respectively instead.

### Do I need a GPU to run Open-AutoGLM?

Not necessarily. The README offers third-party model services from Zhipu BigModel and ModelScope, where you supply a base URL, a model name and an API key. Self-hosting AutoGLM-Phone-9B with vLLM or SGLang is the alternative and does require GPU capacity.

### Why can't Open-AutoGLM type text on my Android phone?

Text entry on Android depends on the ADB Keyboard app, which must be installed and enabled either in the keyboard settings or with adb shell ime enable com.android.adbkeyboard/.AdbIME. HarmonyOS devices use their native input method and do not need it.

## Sources

- [Issues](https://github.com/zai-org/Open-AutoGLM/issues)
- [License: Apache-2.0](https://github.com/zai-org/Open-AutoGLM/blob/main/LICENSE)
- [Project website](https://autoglm.z.ai/blog)
- [README](https://github.com/zai-org/Open-AutoGLM/blob/main/README.md)
- [zai-org/Open-AutoGLM on GitHub](https://github.com/zai-org/Open-AutoGLM)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/zai-org-open-autoglm
