Model or dataset
callstack/agent-device avatar
callstack/agent-device

agent-device: a live app feedback loop for coding agents

Mobile app automation and verification for AI coding agents. CLI, MCP server, and typed Node.js API for iOS, Android, HarmonyOS, TV, web, macOS, and Linux.

4,817 stars315 forksTypeScriptMIT

At a glance

What is it?
Callstack's agent-device gives AI coding agents a way to drive and verify real apps on iOS, Android, HarmonyOS, TV, web, macOS and Linux. It is an MIT-licensed TypeScript project with a CLI, an MCP server and a typed Node.js client.
Who is it for?
Adopt agent-device if your agent already writes mobile code and you need it to check the result in a running app rather than in a diff. Skip it if your test suite is a settled Maestro or Detox setup and nobody is asking an agent to verify anything.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap agent-device fills between a code change and a running app

A coding agent can edit a React Native screen, run the type checker, and report success. It cannot see that the button moved off screen or that the list re-renders on every keystroke. The README frames the project around exactly that gap: giving agents a live app feedback loop so they verify changes in the running app. The audience is teams whose agents already write mobile code and whose review process currently ends at a pull request. The README mentions Expensify using it for mobile bug evidence and profiling, and a Shopify engineer posting about it. The project covers iOS, Android and HarmonyOS on simulators, emulators and physical devices, plus tvOS, Android TV, Amazon Vega OS TV, web, macOS and Linux. That breadth is the point and also the first thing to be sceptical about: a single CLI spanning that many runtimes will not be equally deep on all of them, and the README itself defers per-platform transports and support caveats to the docs.

Snapshots, refs and the settle cycle

The mechanism is accessibility snapshots rather than screenshots. An agent runs a snapshot command, gets back a token-efficient tree of roles, labels and test IDs, and acts on elements by ref. Each action can carry a --settle flag, and the command's diff is the source of truth for what changed. The README is explicit that refs are only valid from the latest output: after a settle command you use the refs in its diff, and you take a new snapshot only when the diff omits what you need. This is a real design constraint, not a footnote. An agent that caches a ref across two actions will press the wrong element, and the failure will look like a UI bug. The README also notes that snapshots come from the accessibility tree, so clear labels, roles and test IDs make runs more reliable, and that screenshots and video are for evidence or for when accessibility data is poor. On top of inspection, the project lists gestures, waits, assertions, alert handling, logs, traces, network data, performance samples, crash details and React profiles, with .ad scripts for replay and Maestro YAML export.

Install agent-device and drive the iOS Contacts app

The README's quick start installs the CLI globally and then asks you to check the environment yourself before handing it to an agent. The documented floor is Node.js 22.12 or newer, with web automation requiring Node.js 24 or newer.

bash
npm install -g agent-device@latest
agent-device doctor
agent-device help workflow

Running doctor is not optional in practice. It reports whether the machine has what each target needs, and the README's own guidance is to run it before the agent does. The help workflow command links to the debugging, replay and profiling guides, and the installed help matches the installed version, which matters when an agent is following instructions you pasted months ago.

The first real session opens an app and inspects it. This example adds a contact in the built-in iOS Contacts app:

bash
agent-device open Contacts --platform ios
agent-device snapshot -i
# @e2 [button] "Add"
agent-device press @e2 --settle
agent-device fill @e7 "Ada" --settle
agent-device screenshot ./contact-form.png
agent-device close

The snapshot with -i returns interactive nodes only, and the README shows @e2 as an Add button. After press with --settle, the diff is expected to include a new text field ref. The README's own output shows the diff adding @e7 for First name, then a later diff showing that field's value changed to Ada and a new Last name field at @e15. Your refs will differ; the README says so directly. Screenshot writes evidence to a path, and close ends the session.

MCP server and the Node.js client are the same execution path

Two integration surfaces sit on top of the CLI. The MCP server starts with agent-device mcp and exposes the installed commands as structured tools over the same execution path as the CLI, so a tool call and a shell call do the same work. Client configuration is a standard stdio entry:

json
{
  "mcpServers": {
    "agent-device": {
      "command": "agent-device",
      "args": ["mcp"]
    }
  }
}

The README points to per-client setup docs and, notably, to guidance on when to prefer the plain CLI over MCP. That is an honest admission that MCP is not always the better interface.

The typed Node.js client is for orchestration code and for agents you build yourself. createAgentDeviceClient takes a session name, and the same commands appear as namespaced methods:

ts
import { createAgentDeviceClient } from 'agent-device';

const client = createAgentDeviceClient({ session: 'qa-run' });
try {
  await client.apps.open({ app: 'com.apple.Preferences', platform: 'ios' });
  const snapshot = await client.capture.snapshot({ interactiveOnly: true });
  const button = snapshot.nodes.find((node) => node.role === 'button');
  if (button) await client.interactions.press({ ref: button.ref });
} finally {
  await client.sessions.close();
}

The package.json exposes subpath exports beyond the root, including ./io, ./artifacts, ./metro, ./batch, ./remote-config, ./android-adb and ./contracts, which suggests the client is meant to be composed rather than used as one monolith. Runnable SDK examples live under examples/sdk in the repository.

Parallel agents, remote devices and where this breaks down

The README states that agent-device coordinates device access across parallel agent worktrees and connects to remote device clouds. That is the feature that decides whether the tool survives contact with a real team. Two agents driving one simulator produce interleaved taps and a snapshot diff that belongs to neither of them, and nothing in the README describes automatic conflict resolution; it describes coordination. The repository has a .worktreeinclude file at the top level, which fits the worktree story, but the README does not document what happens when two sessions target the same device.

The clearer limitation is the accessibility dependency. The README says plainly that screenshots and video are for evidence or for when accessibility data is poor. Apps built with custom drawing, canvas-heavy UI, or unlabelled controls give the agent a thin tree, and the fallback is visual evidence that a model has to interpret. A game or a maps-heavy screen is close to the wrong target. So is a project where the bottleneck is a flaky backend rather than UI state, since agent-device verifies the app, not the service behind it. Finally, the README's platform list is long, and the command docs are where per-target evidence support is defined; assuming a capability exists on HarmonyOS because it works on iOS is a mistake the README warns about indirectly by deferring to those docs.

agent-device vs Maestro and the replay-script question

Maestro is the obvious comparison, and the project invites it by exporting strict Maestro YAML. The difference in approach is who authors the flow. Maestro flows are hand-written declarative YAML that a human maintains and a CI runner executes. agent-device starts from an agent exploring a live app through snapshots and refs, and the README's own prompt example is to explore the checkout flow once, save it as a replay script, and run it in CI. So the .ad script is a byproduct of exploration rather than the starting artifact. That is a genuine advantage when the flow is new and nobody has written the YAML yet. It is a disadvantage when the flow is already written, reviewed and stable, because regenerating it adds a step and a source of drift. The export path exists precisely because teams will want their existing runner to own the final artifact. If your suite is a mature Maestro or Detox setup with no agent in the loop, agent-device adds a second way to do the same thing.

Licence, release cadence and what upgrades cost

The package is MIT licensed, published as agent-device on npm, with the MCP name io.github.callstack/agent-device and the repository at github.com/callstack/agent-device. MIT means you can use it commercially and modify it, with the usual requirement to keep the licence and copyright notice; that is a description of the licence text, not legal advice, and your own counsel should review anything you redistribute.

The release history runs v0.20.9 on 2026-08-17, v0.20.10 on 2026-08-24, and v0.21.0 on 2026-09-08, with package.json at 0.21.3 and the last push on 2026-09-10. That is a fast, pre-1.0 cadence. The practical cost is not the upgrade command, it is that agent-facing surfaces move: tool names, ref semantics and command flags can shift between minor versions, and any prompt or MCP tool list you have pinned to an agent may need re-checking. The README's note that installed help always matches the installed version is the mitigation the project offers, and it is worth leaning on: have the agent read agent-device help workflow from the machine it runs on rather than from a document you wrote earlier. The repository also carries a CHANGELOG.md, which is where a version bump should be checked before upgrading a fleet of CI runners.

Editorial conclusion

Adopt agent-device if your agent already writes mobile code and you need it to check the result in a running app rather than in a diff. Skip it if your test suite is a settled Maestro or Detox setup and nobody is asking an agent to verify anything. Before committing, run agent-device doctor on each target machine, confirm the Node.js version against the documented floor of 22.12 (24 or newer for web), and read docs/commands to see which evidence each platform actually supports.

Frequently asked questions

What is an agent and how does it work?

In this project's context, an agent is a coding assistant such as Claude Code, Codex, Cursor, Windsurf, Cline or Goose that runs a CLI or connects over MCP. agent-device gives that agent a live feedback loop into a running app so it can inspect state, act on visible UI, and save evidence.

What are the risks of using AI agents?

The README does not discuss agent risk in general. It does describe one concrete failure mode: refs are only valid from the latest output, so an agent that reuses a stale ref will act on the wrong element. Running agent-device doctor before handing the CLI to an agent is the project's own stated precaution.

How does an agent router work?

agent-device does not document an agent router. The closest mechanism it describes is the MCP server started with agent-device mcp, which exposes the installed commands as structured tools over the same execution path as the CLI.

What are agent apps?

The README does not define the term. It does say agent-device can be the runtime under agents you build with the AI SDK or Eve, and that the typed Node.js client lets you expose the same commands as model tools in your own agent.

agent device vs maestro

agent-device exports strict Maestro YAML, so the two are complementary rather than mutually exclusive. The difference is authorship: Maestro flows are hand-written declarative YAML, while agent-device has an agent explore a live app through snapshots and refs, then save working steps as a .ad script that can be exported for CI.

Official sources

  1. callstack/agent-device on GitHub
  2. License: MIT
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/callstack-agent-device.svg)](https://hysenlabs.com/projects/callstack-agent-device)