Open-source project
physiclaw/PhysiClaw avatar
physiclaw/PhysiClaw

PhysiClaw reads the screen with a camera and taps it with a stylus

The AI agent that interacts with you in the real world.

384 stars39 forksPythonMIT

At a glance

What is it?
An agent that operates a phone the way a hand would, on the argument that simulated input leaves fingerprints anti-bot systems flag. The cost is seconds per action. The hardware, the CLI and the release cadence are three separate things.
Who is it for?
PhysiClaw fits someone who needs a phone task done on an app with no usable interface for automation, and who is willing to accept a slow, visible, hardware-dependent agent rather than a script. It does not fit unattended batch work, since a dedicated phone, a confirmation pause on payments and seconds per action are all deliberate.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The screen is the interface, and seconds per action is the price

The argument for a physical body is not aesthetic. The applications that run daily life mostly expose no public API, and the usual alternatives to a public API leave marks: desktop automation or Android debugging input produces software fingerprints that anti-bot systems are built to flag.

So the design decision is to treat the screen itself as the interface. A camera reads it, a stylus performs the gestures, and from the phone's point of view the input is indistinguishable from a finger. That is what buys reach, because there is no per-app integration to write and nothing for an app to detect.

The cost is stated plainly rather than buried. A few seconds per action, in exchange for universality and reliability.

That number is the whole design brief. Anything that needs a hundred actions pays a hundred times that latency, which is fine for ordering a coffee and wrong for a data entry job. The target workloads named are errands rather than pipelines: ordering takeout, shopping for groceries, booking a ride, paying a bill, replying to a message.

A dedicated phone with its own messaging account is the front door

The interaction model is deliberately ordinary. The agent has its own phone running its own messaging account, you add that account as a contact, and you write to it the way you would write to anyone.

text
Order a latte on DoorDash.
What's the weather tomorrow?
Buy milk and eggs on Instacart.

The implication for setup is that there is no dashboard to learn and no account to create. The cost is that you carry a second phone, and that this phone is a dedicated device rather than your own, which is what keeps your personal accounts and your agent's actions separate.

That separation is also the security story, or at least the beginning of it. Because the agent holds credentials for one account rather than yours, a mistake touches one phone's session rather than your entire digital life. The project does not make stronger claims than that, and the licence and warnings around payment confirmation suggest the authors are conscious of where the risk sits.

The runtime wakes, works, asks, replies and exits

Execution is event-driven rather than conversational, and the loop is described in five steps.

The runtime runs continuously and wakes the agent in two situations: a scheduled task is due, or the screen lights up with a new message. When it wakes, the agent unlocks the phone, reads the message, and does the task.

Then the part that matters for trust. The agent pauses for confirmation when it matters, and a payment is the example given. It replies with the result, saves its memory, and exits until it is next woken.

Exiting rather than idling is worth noting, since it means the agent is not holding the phone between tasks. A phone left unlocked with a session active is a phone anyone in the room can use, and a loop that closes after each job narrows that window to the length of one task.

The memory save is what makes the next task cheaper, and it is the closest thing here to the persistent context that agent frameworks usually provide. It lives on the agent rather than in a session you would otherwise have to reconstruct.

Hardware and software are released on separate version lines

The repository publishes two things and versions them independently.

One is the agent software, a Python package on its own version line. The other is the hardware, and its releases are tagged with a distinct prefix naming the assembly manual, the sourcing guide and the printed parts.

Three of the most recent releases are all hardware ones, close together in time, which suggests the mechanical design is the part still moving.

The build tooling makes the split explicit. The Makefile carries hardware targets and a variable for the hardware version, with the release tag composed from that value and a note that the version is set per release, and a target that builds and publishes in one command. Release assets are packaged from an output directory inside the hardware tree.

So installing PhysiClaw means matching a software version to a hardware revision, and there is no single version number that covers both. For anyone reproducing a documented setup, that is the detail to record.

More build targets exist for the hardware than for the tests

The Makefile is a good map of how this project is organised, because the target names split cleanly into software and hardware.

On the software side there is the usual set: a fast unit suite as the default, a coverage variant, an alias for the fast suite, targets for slow tests only and integration tests only, everything together, mutation testing, linting, formatting, type checking, and the version bump, build, publish and release chain.

On the hardware side there are about twenty: help, parts, build, check, step, print, a manual, the manual as a PDF, a sourcing document, a drawing, a mark step, a replay step, a camera step, a rebuild, a preflight, deploy, release and an internal packaging target.

Two conventions in that file are worth calling out. The hardware pipeline runs through a Python module with a dependency group needed only for geometry stages, while the manual, sourcing and drawing builders use the standard library alone, which keeps document generation runnable on a machine with no CAD toolchain installed. And the version used by the build and publish targets is scraped out of the package manifest, so a contributor types the version once.

Dependency comments explain what broke and why the pin exists

The package manifest carries comments that are more informative than the version numbers they sit next to, and they read like a changelog of failures.

The model context protocol dependency is bounded below a major version and above the next one, with a comment naming the protocol date, the server module the code imports, the client library handed to the transport, and the single error type the server forwards. It records that the earlier line receives security fixes only, and why the upper bound exists: a later major version broke every import path the earlier one used.

The HTTP clients are two separate packages rather than one. One is imported directly and carries the SOCKS extra so proxy URLs work, with a note that a module rewrites the scheme used by VPN clients on Linux. The other exists because the agent hands the transport a specific client object.

Elsewhere a YAML parser is chosen over the more common one for a specific compatibility reason: the macro files in this project rely on unquoted words being read as strings rather than as booleans.

For a project at alpha status, that density of explanation about why a version is pinned is a reasonable signal about the maintenance burden.

Editorial conclusion

PhysiClaw fits someone who needs a phone task done on an app with no usable interface for automation, and who is willing to accept a slow, visible, hardware-dependent agent rather than a script. It does not fit unattended batch work, since a dedicated phone, a confirmation pause on payments and seconds per action are all deliberate. Two things to check before buying parts: the project is classified alpha and only macOS appears in its package classifiers even though the installer documents Windows and Linux, and hardware releases are tagged separately from the Python package, so pin both if you care about reproducing a setup.

Frequently asked questions

What hardware does PhysiClaw need?

The arm, stylus and camera that make up the device, plus an iPhone for it to operate. The command-line tool itself runs on macOS, Windows and Linux, with a separate installer script for Windows.

Does PhysiClaw install anything on the phone?

No. There are no APIs, no OAuth, no debugging cables and nothing installed on the phone. You unlock it, set it on the desk, and the agent reads the screen with a camera and performs gestures with a stylus.

How do you give PhysiClaw a task?

The agent runs its own dedicated phone with its own messaging account, so you add it as a contact and write plain-language messages such as ordering a latte or buying milk and eggs.

How does PhysiClaw handle a payment?

The agent pauses for confirmation when it matters, with a payment given as the example, then replies with the result and saves its memory before exiting.

Are PhysiClaw hardware releases versioned with the software?

Separately. Hardware assets are tagged under their own release prefix and built from an output directory in the hardware tree, while the Python package carries an independent version.

What trade-off does PhysiClaw accept?

Speed: a few seconds per action, in exchange for universality and reliability, since simulated input leaves software fingerprints that anti-bot systems flag.

Official sources

  1. License: MIT
  2. physiclaw/PhysiClaw on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/physiclaw-physiclaw.svg)](https://hysenlabs.com/projects/physiclaw-physiclaw)