# The bug triage policy is the real documentation

> PokeClaw runs Gemma 4 on the phone and drives the interface from there, so there is no cloud call and no API key in local mode. What makes it worth reading is not the architecture diagram but the two lists in the README that separate harness bugs from model limitations, because that distinction is what most agent projects skip, and it is the reason the project is honest about a 45-second warm-up on CPU-only hardware.

**agents-io/PokeClaw** — PokeClaw (PocketClaw) — first on-device AI that controls your Android phone. Gemma 4, no cloud, no API key. Poke is short for Pocket.

- Repository: https://github.com/agents-io/PokeClaw
- Stars: 1,075 · Forks: 153
- Language: Kotlin
- License: Apache-2.0
- Published: 2026-09-30 · Updated: 2026-09-30 · Language: en
- Canonical page: https://hysenlabs.com/projects/agents-io-pokeclaw

## Phone to model to phone, and what the diagram leaves out

The README contrasts two paths. Everyone else's path is phone to internet to a cloud API and back, with a credit card, an API key and a monthly bill attached. PokeClaw's local path is phone to the language model to phone. In local mode the model runs on the device, and no account or API key is required. The privacy consequence is real for the model call: your prompts and whatever the model reads off your screen do not cross a network to a provider. Two caveats are worth stating, because the diagram is a sales graphic rather than a packet trace. The first is that the phone still talks to the internet for everything else, since reading a WhatsApp conversation means WhatsApp is running, and the model is local while the messages are not. The second is that the app also supports optional cloud models, described as available when you want stronger reasoning for harder tasks, so local-first is a default rather than a guarantee. Nothing in the README says which tasks are considered hard, or what data a cloud fallback would send.

## Reading the screen is the hard part, not the conversation

The author's own framing of what is interesting here is the most useful sentence in the README. The point is not chatting with a local model; the point is getting a local model to read the screen, choose tools, operate apps, keep task state and finish a real phone workflow. That is a fair description of the difficulty, because a chat interface is a text box and a phone agent is a control system. The demo that shows the difference is a good one. In the auto-reply demo the app monitors messages from a contact, reads what was said and replies using the on-device model. In the context demo the contact asks what they were told to bring, and the agent opens the chat, reads the full conversation on screen, finds the earlier message about wine, and answers correctly. The README's own summary of that pair is that it is the difference between context-aware and context-free replies, which is the right benchmark, because an agent that answers from the newest message alone looks identical to a working one until it is wrong.

## Forty-five seconds of warm-up, unless you have a tensor chip

The README explains why one of its demo clips is slow, and this is the single most useful piece of engineering information in the document. The clip was recorded on a CPU-only Android device with no usable GPU or NPU path, and running Gemma 4 E2B on pure CPU takes about 45 seconds to warm up. On better hardware the same model starts in seconds, and the list is specific: Tensor G3 and G4 in the Pixel 8 and Pixel 9, Snapdragon 8 Gen 2 and Gen 3 in the Galaxy S24 and OnePlus 12, Dimensity 9200 and 9300 on recent MediaTek flagships, and Snapdragon 7+ Gen 2 or newer for mid-range devices with a GPU. The summary line is the useful one: same model, better hardware. So the capability question and the hardware question are the same question, and a benchmark from a slow device tells you almost nothing about the experience. The runtime named alongside Gemma is LiteRT-LM, Google's on-device inference library, and the model is described as having native tool calling on it, which is what makes an on-device agent plausible rather than a chatbot with a screen reader.

## Two lists, and the one on the left is the engineering

The README contains the most unusual documentation in this category, a triage policy that separates deterministic harness problems from model limitations. The first list is what should be fixed immediately: the model cannot observe the relevant screen text, a tool reports success when it actually failed, task state leaks into the next run, accessibility, notification access, foreground service or install and update state is wrong, model download, storage or LiteRT startup fails before the model can run, and GPU fallback is mislabelled, crashes, or hides the real reason for the fallback. The second list is what should usually be treated as a model limit: one cloud model fails one multi-step flow once, one local model cannot reason through a complex interface path, a workflow sometimes succeeds because the model chose a weak plan, a prompt needs repeated attempts. The distinction is the point. A tool that lies about success is a bug you can fix and test. A model that picks a weak plan is a fact about the model, and rewriting the prompt for one demo is the failure mode the README explicitly forbids.

## A harness with four layers, and a rule against one-off fixes

The product direction section says the project is becoming a mobile agent harness rather than a chat app with phone control glued on, and then names the four layers. A generic tool layer for phone control. A task and runtime loop that lets a model choose and chain those tools. Playbooks, rules and guards that can be iterated against real device QA. And a product shell on top so the harness is usable by people who are not developers. Then the investment list: generic tools before narrow workflows, repeated real-device QA instead of one-off demos, cloud versus local comparisons on the same task families, and playbooks only when the model proves it needs extra structure. The closing rule is the one to hold the project to. Prompts, tools, skills and playbooks should stay generic, structure should be added when it improves a reusable class of tasks, and one-off prompt hacks or coordinate scripts should be avoided even when they would make a single demo pass. That is an expensive principle to hold, because the demo is what people react to, and it is the difference between a harness and a pile of app flows.

## A different deployment category from the frameworks it cites

The README positions the project against named systems rather than dismissing them, listing DroidRun and Mobilerun, minitap and mobile-use, Mobile-Agent, and AppAgent-style systems as useful references for planning, interface observation, benchmark design and failure recovery. The distinction it draws is deployment, and the first comparison is the clearest: the local open source version of DroidRun and Mobilerun runs the agent loop in a Python command line tool or container on a host machine and talks to the device from there. PokeClaw runs the loop on the phone. That is a structural difference with practical consequences. A host-side loop can be edited and inspected with ordinary development tools, can use a model on a workstation, and can be debugged with a console, but it needs a machine, a link to the device, and permissions to see the screen. An on-device loop needs no companion process, works with no network, and dies with the phone, and it is bounded by what a phone's accelerator can do. For a developer building an agent framework, the first is easier. For an app you actually use, the second is the point.

## A repository that is also a working notebook

The file list is unusual and worth reading as a signal. Alongside the expected app directory, build files and Gradle wrapper there are documents at the root named for the process: an index, a backlog, an execution plan, a strategy file, a quality assurance checklist, a release guide, an architecture reconstruction note, a notice file, and two files for agent instructions. A solo project keeping its planning in the repository rather than in a notes application is defensible, and for a contributor it is a gift, because the plan and the checklist are in the same tree as the code. It also means the repository is partly a working notebook, and a reader has to work out which documents are current. Two more entries deserve a question rather than an assumption. There is a directory of prototypes alongside the app, and a directory of mockups, so the product's history is in the tree rather than in a wiki. And there is an encrypted environment file committed at the root, next to a signatures directory, which on a project that talks about signing and install state is the first thing to ask about before you build from source.

## Prototype framing, and a release line that stopped in May

The README calls the current public build a local-first prototype, and the stated focus is broader device support, more generic skills, more local model options and a cleaner public release path. That last item is a tell, because a cleaner release path implies the current one has rough edges. Builds come from the repository's releases rather than from a store, and the address the README's release badge points at is:

```bash
https://github.com/agents-io/PokeClaw/releases/latest
```

The version line supports the reading: v0.6.12 in April 2026, then v0.7.0 and v0.7.1 on the same day in May, and the last push to main on 2026-06-02, roughly four months before this article. Two releases on one day followed by a quiet branch is the pattern of a project that shipped a version bump and then went quiet, not one that abandoned it. The author asks for real device reports, saying that is how this gets better fast, which is the honest way to grow a project whose hardest problems only appear on hardware you do not own. For an adopter it means the version you install is from a project in progress, on a device that may not be in the list of hardware where the warm-up takes seconds.

## Conclusion

Adopt PokeClaw if you want to experiment with a genuinely on-device phone agent, you have a phone with a recent tensor or GPU path, and you are comfortable installing a prototype build that the project itself calls a local-first prototype. Do not adopt it for unattended automation of conversations on your behalf, because a small local model replying to a family message with the wrong context is a relationship problem rather than a bug, and the README's own context demo is the argument for why. Verify four things first: your hardware against the list that brings warm-up down from 45 seconds, which is Tensor G3 and G4, Snapdragon 8 Gen 2 and 3, Dimensity 9200 and 9300, or Snapdragon 7+ Gen 2, that you need Android 9 or newer, what the encrypted environment file at the repository root contains, since committing one is a thing to ask about, and which build you are installing, given v0.7.1 from 2026-05-26 and a last push on 2026-06-02. The licence is Apache-2.0.

## FAQ

### What does PokeClaw do on the phone?

It runs Gemma 4 on the device through LiteRT-LM and uses it to read the screen, choose tools, operate apps, keep task state and complete phone workflows. The README's example is monitoring a chat, reading the conversation on screen and replying in context, including finding an earlier message in the thread to answer a follow-up question.

### Do I need an API key or a cloud model to use PokeClaw?

No for local mode, which runs the model on the device and requires no account or API key. The app also supports optional cloud models for harder tasks, and the README does not say which tasks qualify or what a cloud fallback would send.

### Why is PokeClaw slow on some phones?

The README says the slow demo clip was recorded on a CPU-only device with no usable GPU or NPU path, and that running Gemma 4 E2B on pure CPU takes about 45 seconds to warm up. Warm-up drops to seconds on Tensor G3 and G4, Snapdragon 8 Gen 2 and 3, Dimensity 9200 and 9300, or Snapdragon 7+ Gen 2 and newer. The minimum Android version is 9.

### How is PokeClaw different from DroidRun or Mobile-Agent?

The README says it is a different deployment category rather than a different implementation. The local open source version of DroidRun and Mobilerun runs the agent loop in a Python command line tool or container on a host machine and talks to the device, while PokeClaw runs the loop on the phone itself. DroidRun, mobile-use, Mobile-Agent and AppAgent-style systems are cited as references for planning, interface observation, benchmark design and failure recovery.

### What licence is PokeClaw released under?

Apache-2.0, with a LICENSE file at the repository root. The most recent release is v0.7.1 from 2026-05-26 and the last push to main was 2026-06-02.

## Sources

- [agents-io/PokeClaw on GitHub](https://github.com/agents-io/PokeClaw)
- [Issues](https://github.com/agents-io/PokeClaw/issues)
- [License: Apache-2.0](https://github.com/agents-io/PokeClaw/blob/main/LICENSE)
- [README](https://github.com/agents-io/PokeClaw/blob/main/README.md)
- [Releases](https://github.com/agents-io/PokeClaw/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/agents-io-pokeclaw
