# understudy watches a demonstration, then replays it as a skill

> Understudy is a local TypeScript agent runtime that drives the GUI, browser, shell and messaging apps of one machine, learns a task from a single demonstration, and generalises it on replay. Its own comparison table marks computer use as macOS only, the npm package sits at 0.3.0 with no GitHub release, and the competitor snapshot is dated March 26, 2026.

**understudy-ai/understudy** — An understudy watches. Then performs.

- Repository: https://github.com/understudy-ai/understudy
- Website: https://understudy-ai.github.io/understudy/
- Stars: 461 · Forks: 36
- Language: TypeScript
- License: MIT
- Published: 2026-09-15 · Updated: 2026-09-15 · Language: en
- Canonical page: https://hysenlabs.com/projects/understudy-ai-understudy

## macOS is the boundary the project states itself

Understudy is pitched as a local agent that operates an entire computer, GUI, browser, shell and messaging, from one instruction, with your own model behind it. An honest limit sits in the project's own comparison table, which lists GUI and computer use for Understudy as yes on macOS only, against a plain yes for two competitors. That matches the demos, which run on macOS with a model reached through Codex and lean on iPhone Mirroring to drive a physical phone. Nothing further down the visible documentation claims another desktop platform, so read the platform row of that table as the specification rather than as marketing.

```text
Day 1:    Watches how things are done
Week 1:   Imitates the process, asks questions
Month 1:  Remembers the routine, does it independently
Month 3:  Finds shortcuts and better ways
Month 6:  Anticipates needs, acts proactively
```

That progression is how the project describes its own layers, from watching to acting ahead of the request. Supporting material sits in the repository too: a product design document, a Chinese README and a Chinese demo page, and a demo directory under `examples/demo-teach/`.

## Teach records intent, and replay swaps the route

Teaching starts with `/teach start` and a single demonstration. What matters is the claim that the skill captures intent rather than coordinates, so it survives a UI redesign, a resized window, or a different application. The published example follows a person: search Google Images for Sam Altman, download a photo, remove the background in Pixelmator Pro, export it, send it to a contact on Telegram, then refine the generated skill interactively before invoking it in natural language. On replay the route can change. A Google Images search becomes browser automation, a download becomes a shell command, while control of a native application stays GUI-driven. The generated artefact is a real skill file, one published under `examples/published-skills/` with a name that spells out the whole task, taught-create-a-background-removed-portrait-for-a-requested-person-and-send-it-in-telegram.

## Eight channels turn a phone into the trigger

Remote dispatch is the demo with the widest blast radius: a message sent from a phone over Telegram is received on the Mac, which converts a file to PDF, opens the desktop Telegram client, finds the right contact and sends it, all through GUI automation rather than an API, with the recording showing the phone and desktop views side by side. Eight built-in destinations are named, Telegram, Discord, Slack, WhatsApp, Signal, LINE, iMessage and the web client, and the comparison table credits Understudy with those eight against twenty or more for one competitor and fifty-plus MCP connectors for another, which are explicitly not a messaging inbox. So the routing works over the apps people already have, and it works by clicking through them.

## A playbook mixes deterministic workers with agentic subagents

The most demanding demo is a six-stage pipeline from one prompt: browse the real App Store in Chrome, install an app on a real iPhone through iPhone Mirroring, explore it autonomously, discover features it was never told about, capture proof-first clips focused on background removal and filters, add English narration and subtitles, compose a vertical video locally with FFmpeg, upload it unlisted to YouTube and clean up the device afterwards, in about an hour with no human intervention. Underneath sits what the project calls workspace artifact composition. A playbook orchestrates two kinds of participant: workers doing deterministic browser and device automation, and skills acting as agentic subagents that make their own decisions. Each stage runs as a separate child session with its own context, and the exploration stage is guided by 51 quality-gate rules while still navigating an app it has never seen.

## Real GUI tests sit behind environment flags

Testing is split three ways, and the split is visible in the scripts. `test` runs vitest, `test:e2e:playbook:synthetic` sets `PLAYBOOK_E2E_MODE=synthetic` for a fake pipeline, and `test:e2e:playbook:live` sets the mode to live. Anything touching a real screen requires `UNDERSTUDY_RUN_REAL_GUI_TESTS=1`, with further flags such as `UNDERSTUDY_RUN_REAL_GUI_GROUNDING_E2E` and a native-task variant narrowing the run. That is the right shape for a desktop agent, since GUI tests are slow and machine-specific and a run that clicks a stranger's screen should be opt-in, and it also means a green default test run tells you nothing about whether the agent can click. Packaging carries the same care: the published `files` list excludes `skills/**/scripts/test_*`, `__pycache__` and `.pyc`, which implies skills ship Python helpers. The suites themselves live in `tests/` and `scripts/e2e/`, and the ordinary `pnpm test` path stays in vitest.

## Default provider is openai-codex with gpt-5.4

Bring your own model, and the shipped defaults are specific:

```bash
UNDERSTUDY_DEFAULT_PROVIDER=openai-codex
UNDERSTUDY_DEFAULT_MODEL=gpt-5.4
```

The example environment also carries empty keys for Anthropic, OpenAI, Google and MiniMax, and one optional value, `UNDERSTUDY_GATEWAY_TOKEN`, described as auth for browser or CLI clients talking to the local gateway. That last one is the piece to think about before exposing anything: the runtime is local-first, but a gateway with a token is a network surface, and an empty token in the example file is a default rather than a recommendation. The npm package is `@understudy-ai/understudy`, an ESM module whose `bin` and `main` are both `understudy.mjs`.

## 0.3.0 on npm, no release, and a March snapshot

Maintenance signals are mixed. `package.json` says 0.3.0 and the repository publishes no GitHub releases, so npm is the only place a version exists and `CHANGELOG.md` is the only history. The last push was on 2026-06-19, and the repository is a pnpm workspace with an oxlint config, a Chinese README and a CLA.md alongside the usual contribution files. More pointed is the comparison table, which declares itself a snapshot as of March 26, 2026 and states that it is conservative, sourced from official docs, with narrow wording where a product does not clearly advertise a capability. Competitor claims age badly, and one row already records a discontinuation on March 25, 2026. Read that table as a positioning document written on a date, not as current market research.

## Where understudy is the wrong tool

Platform decides more than anything else here. Computer use is claimed for macOS only and the marquee pipeline needs iPhone Mirroring, so the most impressive demo is also the least portable. Evidence is thinner than the videos suggest: the claims come with YouTube links, one full unedited recording hosted on Google Drive and a published skill file, which is more than a landing page but is not a test suite result, and the GUI suites that could serve as one sit behind an environment flag. Blast radius is the practical worry. An agent that opens your chat client and sends files on your behalf is doing something you cannot undo from a distance, and the approval modes are described only as far as the visible documentation goes.

## Conclusion

understudy fits a macOS user who wants repeated desktop chores turned into a named skill and triggered from a phone, and who is willing to run an agent against their own logged-in applications. It does not fit Windows or Linux, where the project's own table claims nothing for computer use, and it does not fit anyone who needs a stability record. Verify first the version you install, since 0.3.0 is the only number published and no GitHub release exists, then run the synthetic playbook test before enabling the live one.

## FAQ

### What does Understudy actually do on a computer?

It operates the GUI, browser, shell and messaging apps of one machine from a single instruction, and it can learn a task from one demonstration rather than a written script. You bring your own model, and the local core runs on your hardware.

### Which platforms does Understudy support?

The project's own comparison table marks GUI and computer use as yes on macOS only, and the demos run on macOS with GPT-5.4 through Codex, including one that drives a physical iPhone through iPhone Mirroring. No other desktop platform is claimed.

### Which messaging channels does Understudy support?

Eight built-in channels are listed: Telegram, Discord, Slack, WhatsApp, Signal, LINE, iMessage and the web client. Dispatch works by automating the desktop app rather than by calling a messaging API.

### How do I choose the model Understudy uses?

The example environment ships UNDERSTUDY_DEFAULT_PROVIDER set to openai-codex and UNDERSTUDY_DEFAULT_MODEL set to gpt-5.4, alongside empty keys for Anthropic, OpenAI, Google and MiniMax. UNDERSTUDY_GATEWAY_TOKEN is optional and used as auth for browser or CLI clients of the local gateway.

### How do you run Understudy's end-to-end tests?

The playbook end-to-end script runs in two modes through PLAYBOOK_E2E_MODE, synthetic or live. Tests that drive a real screen need UNDERSTUDY_RUN_REAL_GUI_TESTS=1, with extra flags such as UNDERSTUDY_RUN_REAL_GUI_GROUNDING_E2E to narrow the run further.

## Sources

- [Issues](https://github.com/understudy-ai/understudy/issues)
- [License: MIT](https://github.com/understudy-ai/understudy/blob/main/LICENSE)
- [Project website](https://understudy-ai.github.io/understudy/)
- [README](https://github.com/understudy-ai/understudy/blob/main/README.md)
- [understudy-ai/understudy on GitHub](https://github.com/understudy-ai/understudy)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/understudy-ai-understudy
