Model or dataset
suyiiyii/AutoGLM-GUI avatar
suyiiyii/AutoGLM-GUI

AutoGLM-GUI: a FastAPI control plane for driving Android phones with a language model

A modern web GUI for AutoGLM, making AI automation of Android devices easy - now evolved into your dedicated automation productivity tool.

1,130 stars175 forksPythonApache-2.0

At a glance

What is it?
A 1,100-star Python project that puts a web interface, a cron scheduler, an MCP server and a scrcpy screen feed in front of the AutoGLM phone agent, shipping as a pip package, a set of desktop binaries and a Docker image.
Who is it for?
AutoGLM-GUI is worth a look if your Android automation is already scheduled, repeated or shared between machines, because that is exactly the surface it adds on top of the upstream agent. It is less compelling as a first taste of phone automation, since a plain chat window would have covered most of that and the desktop build already bundles Python and ADB for you.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 10, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What AutoGLM-GUI puts in front of the AutoGLM agent

The upstream project this wraps is a phone agent that takes screenshots, reasons about them and taps. What suyiiyii built around it is everything a person needs in order to keep that agent running: a browser interface, a chat history that survives restarts, a device manager, a scheduler, a built-in MCP server and a live video feed of the handset so you can watch what it is doing.

The package metadata describes it as a Web GUI for AutoGLM Phone Agent with the description AI-powered Android automation, requires Python 3.11 or higher, and carries the Apache-2.0 license. It is a FastAPI application, and the dependency list in `pyproject.toml` tells you most of the design: `apscheduler` for cron style scheduling, `fastmcp` for the MCP server, `openai` and `openai-agents` for the agent loop, `zeroconf` for device discovery over mDNS, `python-socketio` for the push channel, `uvicorn[standard]` for the server, and `prometheus-client` because someone expected to scrape metrics off it. A `droidrun` extra sits in optional dependencies, so that integration is opt-in.

That combination is the interesting part. A cron scheduler inside the same process as the device control means unattended runs are the design centre rather than an afterthought, which is unusual for a GUI project at 1,128 stars.

The interface runs on port 8000, and model credentials are entered in a settings page rather than through environment variables, which is convenient and also means the key lands in the container's config volume at `/root/.config/autoglm`.

The layered agent splits planning from tapping

Two modes ship. The classic mode is a single model doing both the thinking and the acting. The layered mode, which the README describes as stricter, splits the job across two layers: a planning layer that decomposes the task and does multi-step reasoning, and an execution layer that observes the screen and performs the atomic actions.

What makes this more than a diagram is that the planning layer drives the execution layer through tool calls, and those calls and their results are visible in the interface as they happen. So a watch on a twelve-step task can see which step is currently in flight and intervene, rather than waiting for the whole run to fail or wander off. The README frames this as suited to tasks needing multiple rounds of interaction.

There is also immediate interruption, which the README advertises as under one second, along with a device panel, multi-device management with state isolation between handsets, and a workflow feature for saving common tasks as named, editable entries.

For an evaluator this matters because single-model loops fail in a characteristic way: they conflate deciding what to do with doing it, and once the tap lands wrong the model often rationalises the new screen state instead of backtracking. A separate planner gives you a place to observe that failure. Whether the interruption latency claim holds on a busy server is something the issue tracker is a better source for than the README table.

Three install paths, three different maintenance stories

There is a desktop route, a Python route and a container route. The desktop build is a PyInstaller style bundle that ships Python and ADB inside it, distributed as a portable EXE for Windows x64, an arm64 DMG for Apple Silicon, and an AppImage, a deb and a tar.gz for Linux. Auto-update covers the Windows installer, the macOS DMG and the AppImage, and explicitly does not cover the portable EXE or tar.gz.

The Python route is the shortest path if you already have an interpreter. The README calls the pip install the recommended one and offers `uvx` as a no-install alternative that always fetches the newest release:

bash
# 通过 pip 安装(推荐)
pip install autoglm-gui

# 或使用 uvx 免安装运行(需先安装 uv)
uvx autoglm-gui

The `--base-url` flag points at any OpenAI compatible endpoint. The README names Zhipu BigModel, ModelScope, or a self-hosted vLLM or SGLang deployment of `zai-org/AutoGLM-Phone-9B`, which is the path to take if the screenshots should not leave your own hardware.

The container route is a two-file affair. The compose file uses host networking by default and keeps two named volumes, one for config and one for logs:

yaml
    image: ghcr.io/suyiiyii/autoglm-gui:main
    container_name: autoglm-gui

    # Linux: use host network for mDNS/USB support (推荐)
    # macOS/Windows: use port mapping instead (comment out network_mode, uncomment ports)
    network_mode: host

    restart: unless-stopped

The Dockerfile is a two-stage build that compiles a Node 20 frontend with pnpm and then installs the Python package on python:3.14-slim, with `adb` and `curl` pulled in from apt. The health check curls `/api/health` every thirty seconds, which tells you the service does expose a machine-readable health endpoint.

Wireless pairing is the step that decides your deployment shape

The README splits device connection by Android version. On Android 11 and later, pairing is done by scanning a QR code, with no cable at all. On Android 10 and below you must first attach USB and enable wireless debugging, and only then unplug.

The QR path is the interesting engineering constraint. Scanning a pairing code over the local network means mDNS discovery, and the README is direct about the consequence: QR pairing depends on mDNS multicast, which can be restricted inside a Docker bridge network, and it strongly recommends host network mode for full feature support. The compose file encodes the same advice in a comment, offering port mapping as the alternative for macOS and Windows.

The bare `docker run` form shows the volumes and the network flag together:

bash
# 使用 host 网络模式运行(推荐)
docker run -d --network host \
  -v autoglm_config:/root/.config/autoglm \
  -v autoglm_logs:/app/logs \
  ghcr.io/suyiiyii/autoglm-gui:main

For a phone on the same WiFi, the manual path is still available: enable developer options and wireless debugging, note the IP and port, then add the device from the interface. That route skips mDNS, so it survives bridge networking that QR pairing does not.

The compose file also carries a commented USB passthrough line mapping `/dev/bus/usb` for Linux, and describes it as optional. The screen preview is built on scrcpy, and the README says you can click and swipe directly on the live view, with coordinate conversion handled for you, which is how you take over when the agent is stuck.

Scheduled tasks are the reason this is not just another chat wrapper

APScheduler sits in the dependency list with a pin of at least 3.10 and below 4.0, and the README describes a cron style scheduler aimed at repeating work: daily check-ins, periodic checks, anything that used to need a human to remember. Combined with the Docker deployment story, the intended use case is a small server running unattended, controlled from a browser.

The v1.5 feature list frames this as a move from personal assistant to automation hub, and pairs it with conversation history so a failed nightly run can be traced afterwards. Conversation records are saved automatically and the history view lets you replay what happened.

The MCP server is the other half of the unattended story. Built in with `fastmcp`, it lets Claude Desktop or Cursor drive the same device, which means the phone agent becomes a tool inside another agent loop rather than an application you open. The README also carves out an audience for this: if you are an AI agent such as Claude Code, it points you at an `AI_USAGE.md` file with install and API guidance before you start reading the human docs.

Taken together, these three features, scheduler, history and MCP, are what separate this from a thin browser front end. Each one only pays off if you are running the same task more than once, which is also the profile where the Alpha classification starts to matter.

Alpha, bilingual docs, and claims the code cannot confirm

A few things are worth reading carefully before committing time. The package declares Development Status :: 3 - Alpha while carrying 1,128 stars, 174 forks and 35 open issues. Star counts for GUI projects around a popular upstream agent track curiosity more than production use, so treat the classifier as the better signal.

The README is in Chinese and links to `README_EN.md`, which is the English version. Both exist in the repository root, along with `AGENTS.md`, `CLAUDE.md` and `AI_USAGE.md`, which suggests the maintainer works with coding agents on the codebase itself. The recent release notes support that reading: v1.5.20 shipped a rename of the integration test directory to `tests/e2e`, dynamic ports for those services, and a Renovate bump of `zeroconf` marked SECURITY.

The performance numbers in the README are marketing table cells rather than measurements: interruption under one second, and a claim that the architecture separates complex task planning from precise execution. Neither is backed by a benchmark in the repository that anyone can point to. The one thing that is measurable from the outside is task completion rate on your own handsets with your own model endpoint, and that is a weekend of work, not an afternoon.

The honest summary is that this is a well-built shell around someone else's agent, with the scheduling and history features that the upstream project does not have. Evaluate it as an operations tool and the Alpha label is fair. Evaluate it as research and the shell gets in the way.

Editorial conclusion

AutoGLM-GUI is worth a look if your Android automation is already scheduled, repeated or shared between machines, because that is exactly the surface it adds on top of the upstream agent. It is less compelling as a first taste of phone automation, since a plain chat window would have covered most of that and the desktop build already bundles Python and ADB for you. Two details deserve attention before anyone schedules anything against it. The package classifiers themselves say Development Status 3 - Alpha, which is a more honest signal than the star count. And the deployment story forks sharply by platform: wireless QR pairing needs mDNS, mDNS needs host networking, and host networking is not available on Docker Desktop for Mac in the same way it is on Linux. Latest release is v1.5.20 from 2026-08-18 with a last push of 2026-09-19 on `main`, 35 open issues and 174 forks. Try one scheduled task and one interrupt before wiring it into anything that runs unattended.

Frequently asked questions

What is AutoGLM-GUI and how is it different from the original AutoGLM phone agent?

AutoGLM-GUI is a web interface and operations layer for the AutoGLM phone agent, adding a FastAPI backend, conversation history, cron style scheduled tasks, multi-device management, a built-in MCP server and a scrcpy screen preview on top of the upstream agent loop. It is Python, Apache-2.0 licensed and published to PyPI as `autoglm-gui`.

How do I install and start AutoGLM-GUI?

Run `pip install autoglm-gui` and then `autoglm-gui --base-url http://localhost:8080/v1`, or use `uvx autoglm-gui` to run without installing. Open http://localhost:8000 and enter your model API details on the settings page. Desktop builds also exist as an EXE, a DMG and an AppImage.

Which LLM backends does AutoGLM-GUI support?

Any OpenAI compatible endpoint. The README names Zhipu BigModel with the `autoglm-phone` model, ModelScope hosting `ZhipuAI/AutoGLM-Phone-9B`, and a self-hosted vLLM or SGLang deployment of `zai-org/AutoGLM-Phone-9B` for keeping screenshots on your own hardware.

Why does the Docker documentation recommend host network mode?

QR code pairing on Android 11 and later relies on mDNS multicast to discover the device, and mDNS is often restricted inside a Docker bridge network. The README recommends `--network host` on Linux to keep that working. If you pair manually by entering the wireless debugging IP and port instead, port mapping is usually enough.

Can I use AutoGLM-GUI from Claude Desktop or Cursor?

Yes. The project includes a built-in MCP server built on `fastmcp`, which the README lists as a way to integrate with Claude Desktop and Cursor. The repository also ships an `AI_USAGE.md` file with install and API guidance aimed specifically at AI agents driving the project.

Official sources

  1. License: Apache-2.0
  2. Project website
  3. README
  4. Releases
  5. suyiiyii/AutoGLM-GUI on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/suyiiyii-autoglm-gui.svg)](https://hysenlabs.com/projects/suyiiyii-autoglm-gui)