# ARTEMIS: natural-language Android automation from Google's PTE Fusion team

> ARTEMIS turns plain-English instructions into device actions on Android, wraps them in an MCP server for coding assistants, and reports 99%+ completion on AndroidWorld. Here is how it installs, how the two execution profiles differ, and where it stops being the right tool.

**google/artemis** — ARTEMIS turns natural-language instructions into reliable Android automation. It automates end-to-end workflows, captures logs, and integrates seamlessly with AI coding assistants such as Antigravity, Codex, and Claude Code.  It also achieves 99%+ success rate on AndroidWorld Benchmark.  Created by Google's Pixel-Test-Engineering (PTE) Fusion team.

- Repository: https://github.com/google/artemis
- Stars: 10,560 · Forks: 1,046
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/google-artemis

## The problem ARTEMIS solves, and who ends up using it

Android instrumentation is deterministic and brittle in the same breath. A UiAutomator selector breaks when a view id changes, and a human has to re-record the flow. ARTEMIS takes the other route: you describe the workflow in natural language, and the agent looks at the screen, decides the next action, and executes it. The README's own demo is a two-app chain, setting up driving routes in Google Maps and then opening YouTube to play a Coldplay song. That is not a unit test. It is the kind of cross-app task that is tedious to script and easy to describe.

The intended audience is visible in the repository layout rather than in prose. There is an mcp_server/ directory, a packages/artemis-client/ workspace package, a playground/, and a config/artemis.jsonc. The Quick Start section targets people who already have an Android device with USB Debugging enabled and who work inside an AI IDE. The README names Antigravity, Cursor, Claude Code, Codex, Windsurf, VS Code, Cline/Roo and OpenClaw as supported clients for the global MCP configuration and the Artemis Mobile Testing Mindset rules file. If none of those names mean anything to your workflow, this is probably not your tool.

Two execution profiles are offered: flash, described as a reactive observe-and-act loop with asynchronous history summaries, typically 3 to 5 seconds per step, and pro, which checks targets before individual actions and returns blocked actions to the Operator for recovery. The default in .env.example is ARTEMIS_DEFAULT_PROFILE=flash. That split is the single most important design decision to understand before adopting it.

## The observe-and-act loop, the hierarchy backend, and where the coordinates come from

The mechanism the README describes is a loop: observe the device, choose an action, act, repeat. Multimodal targeting is the interesting part. The project says it uses element indices when available, and falls back to coordinate and visual locating for custom interfaces. That ordering matters. Index-based targeting is fast and stable when the UI hierarchy exposes what you need; visual locating is the escape hatch for canvases, games and custom-drawn surfaces where no useful node exists.

Where the hierarchy comes from is configurable. ARTEMIS_HIERARCHY_BACKEND accepts auto, helper or uiautomator, and the comment in .env.example explains the order: auto means the Accessibility Helper first, with UIAutomator2 as the fallback. The dependency list backs this up, with both adbutils and uiautomator2 pinned in pyproject.toml. Choosing helper explicitly means you depend on the Accessibility Helper APK being present on the device. Choosing uiautomator means you skip that install and accept whatever UIAutomator2 can see.

Orchestration is not hand-rolled. pyproject.toml pulls in langgraph, langchain, langchain-core, langchain-community and langchain-mcp-adapters, alongside provider SDKs for Gemini, Vertex AI and OpenAI. So the agent graph is a LangGraph construct, and the MCP layer is bridged through langchain-mcp-adapters. The practical consequence is that model choice is a configuration concern, not a code change, and the README advertises Gemini, Claude, GPT-4o and Qwen-VL as multimodal backends. The trade-off is a deep dependency tree for what is, at the surface, a device automation tool.

## Installing ARTEMIS and running a first task

The Quick Start assumes an Android device or emulator is already connected with USB Debugging enabled. The one-click script is documented as installing ADB, scrcpy, FFmpeg and the Python uv dependencies, then prompting to mount the global MCP server and the testing rules into detected IDEs. Clone and launch on macOS or Linux:

```bash
git clone https://github.com/google/artemis.git && cd artemis
./start.sh
```

On Windows PowerShell the equivalent is the batch file, and the README is explicit that PowerShell does not search the current directory for executables by default:

```powershell
git clone https://github.com/google/artemis.git
cd artemis
.\start.bat
```

The README states that the script opens http://localhost:8000 in your default browser with a device connection wizard, live screen mirroring, a prompt sandbox and execution replays. If you would rather skip the web UI, the README gives a direct CLI invocation with the flash profile:

```bash
uv run artemis run "Open Settings, find Battery and tell me current level" --profile flash
```

Before any of that produces useful output you need a model credential. Copy the template and fill in at least one provider key:

```bash
cp .env.example .env
```

The template lists GEMINI_API_KEY, GOOGLE_API_KEY, OPENAI_API_KEY, ANTHROPIC_API_KEY, OPEN_ROUTER_API_KEY and XAI_API_KEY, with OCR_API_KEY marked optional for Google Cloud Vision OCR. If you want the artemis command available outside the repository, the README suggests uv tool install -e . from the project root. The README does not document what happens when no key is set, so treat that as unverified.

## Wiring ARTEMIS into an IDE over MCP

The MCP server is the feature that distinguishes ARTEMIS from a standalone automation runner. The README describes it as a native Model Context Protocol server that connects a real phone to AI IDEs, and it can be installed non-interactively:

```bash
uv run artemis mcp --install antigravity
uv run artemis mcp --install all
```

The first targets Antigravity (the README also calls it Jetski in the same command comment); the second covers all supported clients including Codex. Interactive setup is available through uv run artemis init. If you prefer to edit configuration yourself, the README points to uv run artemis mcp --generate-config <client> to emit the right TOML or JSON, and shows the Codex shape explicitly:

```toml
[mcp_servers.artemis]
command = "/path/to/artemis/.venv/bin/python"
args = ["-m", "mcp_server"]
cwd = "/path/to/artemis"

[mcp_servers.artemis.env]
PYTHONUNBUFFERED = "1"
PYTHONPATH = "/path/to/artemis"
```

Note the shape of that config: command points at the virtualenv Python, not at a console script, and cwd is the repository root. The Antigravity variant lives at ~/.gemini/jetski/mcp_config.json and uses the mcpServers JSON form. The README's truncated excerpt does not show the full Antigravity payload, so generate it rather than transcribing it. The payoff, per the README, is that the assistant can drive test devices and collect Logcat output and screenshots, which is the loop that turns a chat request into a diagnostic report.

## Where ARTEMIS is the wrong tool

The README's own framing gives the first limitation away. Flash is a reactive loop that acts and then observes the result; pro checks targets before acting and hands blocked actions back to the Operator. Neither is a deterministic assertion engine. If your suite needs to fail a build because a specific view property changed, an LLM-driven loop adds variance where you wanted none. The 99%+ AndroidWorld figure is a benchmark result, and benchmarks are not your app.

Device state is the second constraint. .env.example exposes ARTEMIS_HELPER_AUTO_INSTALL, which the file says lets tasks install or upgrade the Accessibility Helper APK on a device they hold, and instructs setting it to false on shared or personal phones, where only artemis helper install installs it. If you cannot install an APK on the target device, the helper-backed hierarchy path is closed to you and you fall back to uiautomator. Similarly, ARTEMIS_KEEP_DEVICE_AWAKE defaults to true, and the comment warns to set it to false on a personal device whose screen should time out. Both defaults assume a lab phone, not someone's daily driver.

The third constraint is cost and latency that the README quantifies only partially. It states 3 to 5 seconds per step for flash. A multi-step workflow across two apps is therefore minutes of wall clock, plus whatever the model provider charges per observation. There is no release history in the repository to inspect, and the README is silent on retry semantics, rollback of partial workflows, and what happens to an in-flight task when the device disconnects. Those gaps matter more than the benchmark number when you are deciding whether to put this in CI.

## How ARTEMIS differs from scripted UiAutomator suites

The obvious alternative is a conventional Android instrumentation setup built on UIAutomator2 or Espresso, where you write selectors and assertions by hand. The difference in approach is not cosmetic. A scripted suite encodes intent at authoring time: you decide now which node to tap, and the test either finds it or fails. ARTEMIS defers that decision to runtime, resolving targets from the UI hierarchy when indices exist and falling back to coordinates and visual matching when they do not. The scripted suite is reproducible and cheap per run; ARTEMIS is tolerant of UI drift and expensive per run.

That trade-off has a practical shape. A scripted suite can run headless in a container with no model key. ARTEMIS needs a multimodal model, which is why .env.example leads with six provider keys. A scripted suite cannot follow a natural-language instruction like the Maps-and-YouTube demo. ARTEMIS cannot give you a stable pass or fail signal on a single property change without you building that layer on top.

The middle ground is worth naming: use ARTEMIS for exploration and for the long-running stability tests the README mentions under Pro Exploration, then encode whatever you learn as a conventional test. The project ships a Dockerfile that builds an Angular frontend and a Python wheel, installs adb, curl and git, and produces a runtime image, so containerized deployment is contemplated. What the Dockerfile does not do is remove the model dependency.

## Licence, upkeep and what upgrading actually costs

ARTEMIS is Apache-2.0, and every source file in the repository carries the standard Apache header with a 2026 copyright line assigned to Google LLC. For most teams that is the permissive end of the spectrum: you can use it commercially, modify it and redistribute it, provided you keep the notices and state changes. The licence text itself is the authority here, not this summary, and the repository does not ship a NOTICE file in the top-level listing, so if your legal process expects one, that is a question for your counsel rather than a fact I can settle.

The maintenance picture is straightforward. The repository is not archived, and the last push was on 2026-09-10, which is recent. There are no retrieved releases, so there is no versioned changelog to read and no upgrade path documented beyond pulling the branch. Version 1.0 is declared in pyproject.toml with requires-python >=3.12, and the dependency floors are aggressive: langgraph >=1.0.2,<2.0.0, langchain >=1.0.0, mcp >=1.26.0, google-genai >=2.22.0. A uv.lock is committed, which pins the resolved graph for reproducible installs, but it also means upgrades arrive as lockfile diffs you have to review rather than as tagged releases.

Upgrade cost therefore concentrates in two places: the LangChain and LangGraph majors, which change agent orchestration APIs, and the provider SDKs, which change model call signatures. The MCP layer is a third surface, since langchain-mcp-adapters and the mcp package both move. Budget for re-running your workflows after any lockfile bump, because the README does not document compatibility guarantees between these layers.

## Conclusion

Adopt ARTEMIS if you already test Android apps and want an agent loop that can be driven from an IDE through MCP, and if you can supply at least one LLM API key and a device with USB debugging. Do not adopt it if you need a deterministic, assertion-based test suite, if you cannot install the Accessibility Helper on the target phone, or if you are looking for the NASA mission of the same name. Before committing, verify that the profile you pick matches your tolerance for blocked actions, check what ARTEMIS_HELPER_AUTO_INSTALL and ARTEMIS_KEEP_DEVICE_AWAKE should be on your hardware, and confirm the MCP config path your IDE actually reads.

## FAQ

### How do I install ARTEMIS on Android?

You do not install ARTEMIS on the phone. You install it on your workstation, connect an Android device or emulator with USB Debugging enabled, then run ./start.sh on macOS or Linux, or .\start.bat in Windows PowerShell. The script is documented as installing ADB, scrcpy, FFmpeg and the Python uv dependencies, and it may also install the Accessibility Helper APK on the device unless ARTEMIS_HELPER_AUTO_INSTALL is false.

### How do I use ARTEMIS?

The README gives a CLI example: uv run artemis run "Open Settings, find Battery and tell me current level" --profile flash. The one-click script also opens http://localhost:8000 with a device connection wizard, live screen mirroring, a prompt sandbox and execution replays, and the MCP server can be installed into supported AI IDEs with uv run artemis mcp --install all.

### What is the difference between the flash and pro profiles in ARTEMIS?

The README describes flash as a reactive observe-and-act loop with asynchronous history summaries, typically 3 to 5 seconds per step. Pro checks targets before individual actions and returns blocked actions to the Operator for recovery, and is positioned for long-running exploratory and stability tests. The default is set by ARTEMIS_DEFAULT_PROFILE, which .env.example ships as flash.

## Sources

- [google/artemis on GitHub](https://github.com/google/artemis)
- [Issues](https://github.com/google/artemis/issues)
- [License: Apache-2.0](https://github.com/google/artemis/blob/main/LICENSE)
- [README](https://github.com/google/artemis/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/google-artemis
