Model or dataset
joshuayoes/ios-simulator-mcp avatar
joshuayoes/ios-simulator-mcp

ios-simulator-mcp: giving a coding agent hands on the iOS simulator

MCP server for interacting with the iOS simulator

2,184 stars98 forksJavaScriptMIT

At a glance

What is it?
An MCP server that wraps simctl and idb so an agent can tap, type, screenshot and record video inside a running iOS Simulator.
Who is it for?
ios-simulator-mcp is at its best as a loop closer: an agent edits code, launches the app, describes the accessibility tree, taps the thing it just changed, and looks at a screenshot. The tools are narrow and legible, every one of them maps to a command an iOS developer already knows, and the udid parameter means multi-device runs do not need special handling.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 54 days ago.
What is it written in?
Mainly JavaScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.

Editorial analysis

A tool per simulator verb rather than one catch-all command

The design choice that makes this server easy to reason about is that each simulator operation gets its own tool with its own parameters. There is no generic shell tool and no way to pass an arbitrary command string. `get_booted_sim_id` returns the identifier of the running simulator, `open_simulator` launches the Simulator application, and `ui_tap` takes coordinates. An agent that wants to press a button needs to know the coordinates, which means it has to look at the screen first, and the tool that lets it look is `ui_view`.

That ordering is the whole point. `ui_view` returns a compressed screenshot as image content, `ui_describe_all` returns the accessibility information for the entire screen, and `ui_describe_point` returns just the element at a pair of coordinates. A client can combine a describe call with a tap call and get a closed loop without the server having to guess what the user meant. The 2.0.0 release notes describe the tool set as fourteen entries whose names, schemas and behaviour were all preserved across the SDK migration, which is a useful number to keep in mind: the tool list in the README walks from `get_booted_sim_id` through `open_simulator`, the `ui_*` family, `ui_view`, `screenshot` and `record_video`, and the release notes count more than the README spells out parameter by parameter.

The `ui_` prefix is doing real work. Everything under it is about the screen contents and its accessibility tree, everything without it is about the simulator as a device. That split tells you where the author's attention went, and it is not evenly distributed. The screen side has nine tools; the device side has the boot lookup, the open call, the screenshot call and the video call.

The accessibility tree is the payload, not a debugging extra

Most automation servers treat accessibility data as metadata. Here it is one of the main outputs, and the tool set reflects that choice. `ui_describe_all` returns accessibility information for the whole screen, which for an agent means a list of labelled elements with their types and frames. `ui_describe_point` is the narrower version: give it a coordinate and it tells you what is there.

`ui_find_element` is the tool added in 1.6.0, and its parameters show what the author was trying to fix. You pass an array of search strings, and an element matches if any string matches against its AXLabel or AXUniqueId. You can narrow by element type, choosing substring or exact matching, and you can turn case sensitivity on. The release note for that version is explicit about the motivation: returning only matching elements instead of the full tree dramatically reduces token usage in agent loops. For a project whose entire consumer is a language model, that is the correct thing to optimise.

The type parameter takes a string such as Button, StaticText or Group, matched case-insensitively and exactly rather than by substring. That is a slightly unusual choice, since an agent that guesses the element type will get nothing back rather than something close, but it does keep the results predictable when the same label appears on several elements of different kinds.

One parameter type repeats across every tool

Every tool that touches a device takes an optional `udid` string, and the schema comment attached to it is the same everywhere: it identifies the target, it can also be set with the IDB_UDID environment variable, and the format is a UUID written as 8-4-4-4-12 hexadecimal characters.

typescript
{
  /**
   * Udid of target, can also be set with the IDB_UDID env var
   * Format: UUID (8-4-4-4-12 hexadecimal characters)
   */
  udid?: string;
}

That single optional parameter is what makes multi-simulator work possible. Without it a server has to assume there is one device, which is true on a clean developer laptop and false the moment a second simulator is running. The 1.6.0 release notes record the same problem from the other direction, when `record_video` stopped hardcoding the string `booted` and gained a `udid` parameter so recording could target a specific device while others were running.

The swiping tool shows how much the schemas spell out. Coordinates are named by their role, `x_start`, `y_start`, `x_end`, `y_end`, and each one carries a comment saying which end it is. Durations are strings that accept decimal numbers. There is an optional `delta` for the size of each step in the swipe, defaulting to 1. For a tool an LLM has to call without reading documentation, having every parameter carry its own units and default is worth more than any amount of prose about it.

Screenshots and video both land in a directory you configure

Two tools produce files. `screenshot` takes an `output_path` and an optional image type, where the format list is png, tiff, bmp, gif or jpeg and the default is png. It also takes `display` to pick between internal and external, with the default depending on device type, and `mask` to choose how non-rectangular displays are handled, with the options ignored, alpha or black.

The path handling is the part worth internalising. If the path is relative, it resolves against the directory named by the IOS_SIMULATOR_MCP_DEFAULT_OUTPUT_DIR environment variable, and if that variable is not set, the file goes to `~/Downloads`. `record_video` follows the same rule and treats its output path as optional, falling back to a generated default name.

bash

The video tool passes a codec option of h264 or hevc, with hevc as the default, and it has a `force` flag to overwrite an existing file. Both tools describe their simulator interaction honestly: the video notes say it records using simctl directly, which is a hint that this path shells out rather than going through the accessibility path the other tools use.

The text input tool is the one with a hard limitation stated in its schema: `ui_type` documents its text parameter as ASCII printable characters only. That is not a stylistic choice, it is what the underlying typing mechanism supports, and an agent trying to enter a password with an emoji will get a partial string.

The 2.0 migration moved the install floor to Node 20

Version 2.0.0 shipped on 2026-08-11 and was almost entirely a dependency migration. The release notes call it a pure migration, stating that all 14 tools keep identical names, schemas and behaviour, so no MCP client configuration changes are needed. The dependency moved from `@modelcontextprotocol/sdk` 1.x to `@modelcontextprotocol/server` at version 2, with zod 4 alongside it.

json
"engines": {
  "node": ">=20"
},

The one user-visible change is that Node 18 support is gone, because the SDK v2 drops it. The package manifest makes that concrete with an engines field requiring Node 20 or newer. Two days later came 2.1.0, which added the app lifecycle tools: `terminate_app` by bundle identifier, `open_url` for https pages, custom URL schemes, universal links and OAuth redirect flows, and `list_apps` returning sorted name and bundle identifier lines so an agent can discover identifiers before launching or terminating anything.

The `list_apps` implementation detail is a good indicator of the project's level of care. Rather than parsing the output of `simctl listapps` directly, it converts the plist through `plutil -convert json`, which the notes describe as resilient to formatting differences across Xcode versions. The same release added an optional `input` option to the shared command helper so data can be piped to a child's stdin without a shell, and that helper change is what `list_apps` uses.

A command injection notice that never came down

The first thing under the title is a security notice, and it is unusually direct: command injection vulnerabilities present in versions below 1.3.3 have been fixed, so update to 1.3.3 or later, with details in SECURITY.md. The repository tree carries that file alongside TESTING.md, TROUBLESHOOTING.md, QA.md, CLAUDE.md and CONTEXT.md, which is a sign of a project that keeps its working notes in version control.

The notice matters more than it might look. A server whose entire job is turning structured tool calls into subprocess invocations against simctl and idb is exactly the shape of program where arguments can end up on a shell command line, and the fix landed at 1.3.3 while the current version is 2.1.0. Anyone pinning an older copy for compatibility reasons should read SECURITY.md first, and anyone writing a policy around which MCP servers are allowed in a development environment should notice that this one once had the bug class.

The rest of the README is a feature list and a screenshot, followed by a Featured In section naming Anthropic's Claude Code best practices article, React Native Newsletter issue 187, a mobile automation newsletter, and the punkpeye awesome-mcp-servers collection. There is no architecture document, no contributing guide on the visible page and no description of what happens when no simulator is booted. The repository is 2,184 stars with 98 forks and 28 open issues, and the last push was on 2026-08-13, the same day 2.1.0 was published.

Editorial conclusion

ios-simulator-mcp is at its best as a loop closer: an agent edits code, launches the app, describes the accessibility tree, taps the thing it just changed, and looks at a screenshot. The tools are narrow and legible, every one of them maps to a command an iOS developer already knows, and the udid parameter means multi-device runs do not need special handling. What the repository does not settle is how reliable each visual interaction is across the long tail of apps, and the README's tool list runs from get_booted_sim_id to record_video without the argument schema for the app lifecycle calls the 2.1.0 notes describe. Start with the Cursor deep-link install, pin at or above 1.3.3 because of the command injection fix, and read SECURITY.md before pointing any agent at a device.

Frequently asked questions

Can Claude Code control an iOS simulator?

Yes, through this server. It exposes simulator operations as MCP tools, and any MCP client can register it, including the clients Anthropic ships. The README notes that the project is featured in Anthropic's Claude Code best practices article, which discusses writing code, taking a screenshot of the result and iterating.

Which simulators can ios-simulator-mcp target?

Any booted iOS simulator. Every tool takes an optional udid parameter, and setting the IDB_UDID environment variable provides a default so the parameter can be omitted. The udid is a UUID in 8-4-4-4-12 hexadecimal form, and passing it explicitly is how you address one device when several are running.

Does ios-simulator-mcp require a jailbroken device?

No. It drives the standard iOS Simulator through simctl and idb, both of which are part of the Xcode toolchain. It does require Xcode to be installed on the machine running the server, since there is no simulator to talk to otherwise. The one environment variable worth knowing about is IOS_SIMULATOR_MCP_DEFAULT_OUTPUT_DIR, which decides where relative screenshot and video paths land.

What changed between ios-simulator-mcp 2.0 and 2.1?

2.0.0 was a dependency migration to MCP TypeScript SDK v2 with zod 4, and its only user-visible effect was dropping Node 18 support in favour of Node 20 or newer. 2.1.0 added the app lifecycle tools: terminate_app, open_url and list_apps, where list_apps converts the simctl listapps plist to JSON before parsing it.

Official sources

  1. Issues
  2. joshuayoes/ios-simulator-mcp on GitHub
  3. License: MIT
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/joshuayoes-ios-simulator-mcp.svg)](https://hysenlabs.com/projects/joshuayoes-ios-simulator-mcp)