mobile-mcp: MCP Server for iOS and Android Automation via Accessibility Snapshots
Model Context Protocol Server for Mobile Automation and Scraping (iOS, Android, Emulators, Simulators and Real Devices)
At a glance
- What is it?
- mobile-mcp is an MCP server that gives AI agents structured access to iOS simulators, Android emulators, and real connected devices through the native accessibility tree. It eliminates the need for platform-specific automation knowledge and works with any MCP-compatible agent client.
- Who is it for?
- mobile-mcp is the right choice for teams building agent-driven mobile automation workflows where the agent controls the device rather than a scripted test runner. It is not the right fit for traditional CI test pipelines where deterministic YAML flows are preferred: tools like Maestro or Appium fit that use case better.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What mobile-mcp Does and Who It Is For
Model Context Protocol (MCP) is a protocol for giving AI agents access to external tools and data sources through a standardized interface. mobile-mcp implements this protocol for mobile devices: it exposes iOS simulators, Android emulators, and physically connected iOS and Android devices as tools that any MCP-compatible agent can drive.
The primary use cases the README lists are: native app automation for testing or data-entry scenarios, scripted multi-step user flows without manually controlling emulators, automating user journeys driven by an LLM, and general-purpose mobile application interaction for agent-based frameworks. The server also supports agent-to-agent communication for mobile automation data extraction.
The target user is an engineer or team building an AI agent that needs to interact with a mobile app as part of a larger workflow. This differs from traditional mobile test automation, where a test engineer scripts explicit tap sequences in advance. mobile-mcp is designed for the pattern where the agent decides what to do at each step based on what it sees on the screen.
The server is compatible with Claude Code, Codex, Gemini, GitHub Copilot, and any other MCP-compatible client, according to the README.
The Accessibility-First Architecture: No Vision Model Required
The most architecturally significant choice in mobile-mcp is its use of the native accessibility tree as the primary interaction model rather than computer vision or screenshot analysis.
On iOS, the accessibility tree is the same data structure that VoiceOver uses. On Android, it is the accessibility framework that TalkBack relies on. Both provide structured information about every visible UI element: its type (button, text field, label), its text content, its position on screen, and its interactability. Reading this tree is fast, deterministic, and requires no image tokens.
When the accessibility tree cannot resolve a tap target, the server falls back to screenshot-plus-coordinates. The README describes this as falling back to screenshots only when needed. In practice this means most agent interactions use the cheaper accessibility path, with vision-based fallback available when the app does not expose its UI through the accessibility API.
The mobile_list_elements_on_screen tool returns UI elements with their coordinates and properties from the accessibility tree. The mobile_take_screenshot tool captures the screen as an image. Using both together, an agent can see the screen visually and also get structured element data, which reduces ambiguity compared to screenshot-only approaches.
Installing mobile-mcp and Connecting Devices
The server is published to npm under the package name @mobilenext/mobile-mcp. The quickest way to run it is via npx:
npx @mobilenext/mobile-mcpThe README documents Xcode command line tools as a prerequisite for iOS targets and Android Platform Tools (adb) as a prerequisite for Android targets. Without these, the server starts but cannot discover or connect to devices.
For VS Code, the README includes a one-click install badge that configures the MCP server in the VS Code insiders build without manual JSON editing. The VS Code configuration passes the command npx with arguments -y and @mobilenext/mobile-mcp@latest.
Device connection follows platform conventions. iOS simulators require a booted simulator visible to xcrun simctl. iOS real devices must be connected over USB and trusted on the device. Android emulators require an Android SDK installation and a running emulator visible to adb. Android real devices require USB debugging enabled and authorized through the Android developer options. The mobile_list_available_devices tool lists all discovered devices once the prerequisites are satisfied.
For cloud devices, the README describes a cloud provider authentication flow using mobile_login_to_cloud_provider, followed by mobile_list_remote_devices and mobile_allocate_remote_device to reserve a physical device from the Mobile Next cloud fleet.
Available MCP Tools: Device Management, Screen Interaction, and Logs
The server exposes over 30 MCP tools across several categories. Device management tools include mobile_list_available_devices, mobile_get_screen_size, mobile_get_orientation, mobile_set_orientation, mobile_set_location (for GPS override), and mobile_clipboard (to read or write the device clipboard).
App management tools cover the full lifecycle: mobile_list_apps lists installed apps, mobile_launch_app and mobile_terminate_app start and stop apps by bundle ID or package name, and mobile_install_app and mobile_uninstall_app handle .apk, .ipa, .app, and .zip files.
Screen interaction tools include mobile_take_screenshot, mobile_save_screenshot, mobile_list_elements_on_screen, and coordinate-based interaction tools: mobile_click_on_screen_at_coordinates, mobile_double_tap_on_screen, mobile_long_press_on_screen_at_coordinates, and mobile_swipe_on_screen. Screen recording is handled by mobile_start_screen_recording and mobile_stop_screen_recording.
Input and navigation tools include mobile_type_keys for text input with optional submit, mobile_press_button for hardware buttons (HOME, BACK, VOLUME_UP, VOLUME_DOWN, ENTER), and mobile_open_url for deep links or browser URLs.
Logs and crash reports are accessible through mobile_get_device_logs (logcat on Android, unified log on iOS), mobile_list_crashes, and mobile_get_crash. The mobile_batch_commands tool runs multiple tools in sequence in a single call, which is useful for common patterns like click-type-click workflows.
Docker Deployment and ADB Configuration
The repository includes a Dockerfile for containerized deployment. The production stage is built on node:22-slim and installs adb for Android device connectivity. The key environment variable for Docker deployments is ADB_SERVER_SOCKET, which tells the containerized server how to reach the host's adb server:
ENV ADB_SERVER_SOCKET=tcp:host.docker.internal:5037This setting connects the container's adb client to the adb server running on the host machine at port 5037, which is the standard adb server port. The container does not run its own adb server; it communicates with the host's server, which must be started separately. This design means device authorization happens on the host rather than inside the container.
The Dockerfile uses a non-root user (node) for the runtime process and marks /app as writable by node, because mobile_save_screenshot writes files below the working directory. This detail matters if you mount a volume for screenshot persistence.
The build step uses --ignore-scripts on npm ci to skip the husky prepare hook, which is a devDependency that should not run in a production image.
Limitations: Platform Prerequisites, Simulator Constraints, and Cloud Cost
The server has hard dependencies on platform tooling. iOS automation requires Xcode command line tools installed on the host. Android automation requires the Android SDK and adb. These are not lightweight prerequisites: Xcode command line tools alone require several gigabytes. For teams without an existing Apple or Android development environment, the setup cost is real.
Real iOS device automation requires that the device be physically connected via USB and trusted. Remote iOS device automation over the network is not supported without the Mobile Next cloud product. For teams that need to test on real devices at scale without a device lab, the cloud option adds a cost dependency.
Apple's sandbox model limits what automation can do on real iOS devices: access to sensitive APIs (camera, location, health data) requires explicit permission grants, and some system dialogs cannot be reliably automated. The README does not document these constraints per-tool, so testing your specific app's permissions is necessary before assuming full automation coverage.
The server's Apache-2.0 license is permissive for commercial use. However, any dependency on the Mobile Next cloud service for remote devices introduces a third-party service dependency that is outside the control of the open-source package.
mobile-mcp vs Appium and Maestro
Appium is the most established mobile automation framework. It uses the WebDriver protocol, supports iOS and Android, and provides a large ecosystem of language bindings and integrations. Appium requires a running Appium server process, a client library, and device-specific drivers (xcuitest-driver for iOS, uiautomator2 for Android). The setup is complex but mature and well-documented.
mobile-mcp differs in purpose: it is an MCP server for AI agents, not a test framework. The tools it exposes are designed to be called by an agent that decides what to do next, rather than by a pre-scripted test. Where Appium tests are deterministic sequences, mobile-mcp agent sessions are exploratory.
Maestro is a YAML-based mobile testing tool focused on scripted flows with minimal setup. It targets the same device types as mobile-mcp but uses declarative YAML flows rather than an agent. The README mentions Maestro as a search comparison term for users evaluating alternatives. For teams that want predictable scripted test coverage, Maestro's model is more appropriate. For teams building agents that decide their next action from screen state, mobile-mcp is the relevant tool.
Recent releases: the project published 1.0.4 on 2026-09-13 and 1.0.3 on 2026-09-08, indicating active release cadence.
Editorial conclusion
mobile-mcp is the right choice for teams building agent-driven mobile automation workflows where the agent controls the device rather than a scripted test runner. It is not the right fit for traditional CI test pipelines where deterministic YAML flows are preferred: tools like Maestro or Appium fit that use case better. Before deploying, verify that Xcode command line tools are installed for iOS targets and that adb is in the PATH for Android targets; the README lists both as prerequisites, and the server will not connect to devices without them.
Frequently asked questions
How do I use mobile-mcp?
Add the MCP server to your agent client's configuration or run it directly with npx @mobilenext/mobile-mcp. The server then exposes tools like mobile_list_available_devices and mobile_take_screenshot to any connected agent. iOS requires Xcode command line tools; Android requires adb and the Android SDK.
What is mobile-mcp?
mobile-mcp is an MCP (Model Context Protocol) server that gives AI agents structured access to iOS simulators, Android emulators, and real connected devices. It drives apps through the native accessibility tree rather than relying on image tokens, falling back to screenshots only when needed.
How does mobile-mcp compare to Maestro?
Maestro is a YAML-based mobile testing tool for scripted deterministic flows. mobile-mcp is an MCP server designed for AI agents that decide their next action from the current screen state. The README does not document a direct feature comparison, but the use cases differ: mobile-mcp is for agent-driven automation, while Maestro is for predictable scripted test sequences.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/mobile-next-mobile-mcp)