minitap-ai/mobile-use: driving Android and iOS apps with a LangGraph agent
AI agents can now use real Android and iOS apps, just like a human.
At a glance
- What is it?
- mobile-use is an Apache-2.0 Python agent that controls a real phone through natural language, built on LangGraph with ADB, uiautomator2, idb and Appium underneath. It is aimed at engineers who need scripted UI work on a device, not at people looking for a phone assistant.
- Who is it for?
- Adopt mobile-use if you have a device or emulator, an LLM API key, and a task whose output you can describe in a schema, such as pulling the first three unread Gmail sender and subject pairs. Do not adopt it for games, since the README states games do not expose accessibility tree data and effectiveness is limited there, and do not adopt it if you need a stable scripting API rather than an LLM in the loop.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 17 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What mobile-use actually automates
Most UI automation on a phone means writing selectors against a resource id or an accessibility label, then re-writing them when the app ships a new build. mobile-use takes the other route: you describe the task in a sentence, and a LangGraph agent reads the current screen and decides the next tap or swipe. The README's own example is "Open Gmail, find first 3 unread emails, and list their sender and subject line", paired with an output description that asks for a JSON list of objects with sender and subject keys.
The audience follows from that. This is for engineers who need a device to do something once or repeatedly without hand-writing the interaction graph: QA teams running a flow across an app, data people pulling structured records out of an app that has no API, and researchers who want an agent loop over a real touchscreen. It is not a consumer assistant, and the README does not present it as one. The project targets Python 3.12 or newer and its package name on PyPI is minitap-mobile-use.
The LangGraph loop and the device layer beneath it
The pyproject description calls mobile-use a "multi-agent system that automates real Android and iOS devices through low-level control using LangGraph". The dependency list shows how that splits. langgraph, langchain, langchain-core and the provider packages (langchain-openai, langchain-anthropic, langchain-google-genai, langchain-google-vertexai, langchain-azure-ai, langchain-cerebras) form the reasoning side. adbutils, uiautomator2, fb-idb, facebook-wda and Appium-Python-Client form the device side, one stack per platform.
Model choice is configuration, not code. The repository ships llm-config.defaults.jsonc with named presets and llm-config.override.template.jsonc to copy over it, and the README says you can set "provider": "minimax" or "provider": "anthropic" on individual agent nodes, or point OPENAI_BASE_URL at any OpenAI-compatible endpoint. That per-node granularity is the interesting part: you can run a cheap model on one node and a stronger one where the screen reading matters. The README lists MiniMax-M2.7 and MiniMax-M2.7-highspeed at 200K context, and claude-sonnet-4-6 with claude-haiku-4-5-20251001, also 200K. Those are the documented options, not a ranking.
Installing mobile-use and running a first Gmail task
There are two documented paths. The local route starts from the repository, where you copy the environment template and fill in provider keys. The file lists MINITAP_API_KEY, ANTHROPIC_API_KEY, OPENAI_API_KEY, GOOGLE_API_KEY, XAI_API_KEY, OPEN_ROUTER_API_KEY and MINIMAX_API_KEY, with OPENAI_BASE_URL commented out for overriding the default endpoint.
cp .env.example .env
cp llm-config.override.template.jsonc llm-config.override.jsoncAfter editing those two files, the installed console script is mobile-use, which pyproject maps to minitap.mobile_use.main:cli. The README points to the local quickstart at docs.minitap.ai/mobile-use-sdk/quickstart for the exact invocation, so treat the CLI flags as something to confirm there rather than guessing them.
The Docker quickstart is the more concrete one, and the README states it is Android-only for now and requires Docker installed. You either plug in an Android device with USB debugging enabled in Developer Options, or start an emulator. On Linux or macOS the wrapper script takes the task and an output description:
chmod +x mobile-use.sh
bash ./mobile-use.sh \
"Open Gmail, find first 3 unread emails, and list their sender and subject line" \
--output-description "A JSON list of objects, each with 'sender' and 'subject' keys"On Windows the equivalent is a PowerShell script invoked with an execution policy bypass. The README warns that if you use your own device you must accept the ADB connection prompts that appear on screen, and that the script connects over IP, so the phone has to be on the same Wi-Fi network as your computer. If it fails with "Could not get device IP. Is a device connected via USB and on the same Wi-Fi network?", the troubleshooting section says it could not find one of the common Wi-Fi interfaces. The docker-compose file exposes the same idea as services: mobile-use-full-usb passes /dev/bus/usb with privileged: true, while mobile-use-full-ip uses ADB_CONNECT_ADDR. Results and events land at the paths set by RESULTS_OUTPUT_PATH and EVENTS_OUTPUT_PATH.
Where mobile-use breaks down
The README is unusually direct about one limit: UI-aware automation "currently has limited effectiveness with games as they don't provide accessibility tree data". That is not a tuning problem. If the app draws to a surface and exposes nothing to the accessibility layer, the agent has little to read, and any task that depends on fast reaction to moving graphics is outside what this design can do.
Three other constraints are visible in the repository rather than stated as warnings. First, the Docker quickstart is Android-only, so the iOS path through fb-idb, facebook-wda and Appium-Python-Client has no equivalent documented one-liner. Second, the IP-based connection model means a device on a different network than your machine is a setup problem, not a runtime one. Third, and most important for anyone planning to put this in CI: an LLM sits in the decision loop. Two runs of the same sentence can take different paths through the UI. If you need a deterministic replay of a fixed tap sequence, a plain uiautomator2 or Appium script is the better tool, and mobile-use is the wrong layer.
mobile-use compared with writing Appium or uiautomator2 directly
The closest alternative is not another agent, it is the libraries mobile-use already depends on. A hand-written uiautomator2 or Appium script names each element and each step. It runs the same way every time, it is fast, and when the app changes you get a clear failure at the line that broke. Its cost is authoring and maintenance: every new flow is new code, and a flow that spans several apps means several sets of selectors.
mobile-use inverts that trade. You write a sentence and an output description instead of a step list, and the agent adapts when a button moves. You pay for it with nondeterminism, token cost on every run, and a dependency on screen data the app may not expose. The honest split is that Appium and uiautomator2 are for flows you will run thousands of times and can afford to maintain, while mobile-use fits exploratory or low-volume tasks where writing the selector graph would take longer than the task is worth. There is also a middle position: the repository exposes an MCP server, documented at docs.minitap.ai/v2/mcp-server/introduction, which lets a host agent call into this device control rather than driving it from the CLI.
Maintenance, licence and what an upgrade costs
The repository is not archived, and its last push was on 2026-09-10, which is recent. Releases, though, are sparse: v2.6.0 on 2025-10-20, v2.9.0 on 2025-11-15, and v3.3.0 on 2026-01-12. Meanwhile pyproject.toml already declares version 4.0.0, so the tagged releases trail the code on main. If you pin to a release you are pinning to something months behind the branch, and if you track main you are tracking code that has moved past its last tag. Decide which of those you want before you build a pipeline on it.
Upgrade cost concentrates in two files. llm-config.defaults.jsonc holds the presets, and the README instructs you to copy a preset object into your own llm-config.override.jsonc, which means a provider change in defaults does not reach you automatically. The override file is also mounted into the containers by docker-compose, so a schema change there breaks the container start rather than a single command. The dependency set is broad by design, spanning several LangChain provider packages and three device stacks, and the Dockerfile pins Python through uv with uv sync --locked, so lockfile churn is the other place upgrades will bite.
Licensing is Apache-2.0, declared in pyproject as license = { file = "LICENSE" } and shown as a badge in the README, with a NOTICE file at the repository root. Apache-2.0 includes an explicit patent grant and requires that NOTICE and attribution be preserved when you redistribute. What that means for your product is a question for your own counsel, not something this article can settle.
Editorial conclusion
Adopt mobile-use if you have a device or emulator, an LLM API key, and a task whose output you can describe in a schema, such as pulling the first three unread Gmail sender and subject pairs. Do not adopt it for games, since the README states games do not expose accessibility tree data and effectiveness is limited there, and do not adopt it if you need a stable scripting API rather than an LLM in the loop. Before committing, verify that your device is reachable on the same Wi-Fi network as your machine, that your chosen provider key works against the agent nodes in llm-config.override.jsonc, and that the iOS path, which depends on fb-idb, facebook-wda and Appium-Python-Client, behaves on your hardware, because the Docker quickstart covers Android only.
Frequently asked questions
How do I install mobile-use and run it for the first time?
Copy .env.example to .env and fill in a provider key, optionally copy llm-config.override.template.jsonc to llm-config.override.jsonc to choose models, then run the mobile-use.sh wrapper with a task sentence and an --output-description. The README states the Docker quickstart currently works for Android devices and emulators only, and that the device must be on the same Wi-Fi network as your computer.
What is mobile-use used for?
It controls a real Android or iOS device from a natural language instruction, for example opening Gmail and listing the sender and subject of the first three unread emails. The README also lists data scraping, where information is extracted from an app into a format such as JSON described in natural language.
Does mobile-use work with games?
The README states that UI-aware automation currently has limited effectiveness with games, because games do not provide accessibility tree data. The agent reads the interface through that layer, so an app that exposes nothing to it gives the agent little to work with.
Which LLM providers can power mobile-use?
The README names OpenAI, Google, xAI, OpenRouter and MiniMax, and the dependency list adds Anthropic, Google Vertex AI, Azure AI and Cerebras. You can also point OPENAI_BASE_URL at any OpenAI-compatible endpoint, including a local model, and set the provider per agent node in llm-config.override.jsonc.
Does mobile-use run on iOS as well as Android?
The dependencies include fb-idb, facebook-wda and Appium-Python-Client for the iOS side, so the code supports it. The README's Docker quickstart, however, is marked as available for Android devices and emulators only, so the documented one-command path is Android.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/minitap-ai-mobile-use)