# typesafe-computer-use gets its 155x cost edge from free output tokens, not a smaller context

> jev, the short name for this Mac automation loop, reads the screen with the operating system's own OCR and accessibility tree, asks a small decision model which of a short list of actions comes next, and only calls a writing model when a text field genuinely needs typing. The cost table is the interesting artifact: the two systems send almost the same number of input tokens, and the difference is entirely on the output side. The rest of the project is the price of that trade, namely rebuilding as deterministic code the reasoning a frontier model does implicitly.

**awlevin/typesafe-computer-use** — Computer use for about $0.0002 a step: OCR the screen, classify the next action with TypeSafe, click. macOS.

- Repository: https://github.com/awlevin/typesafe-computer-use
- Stars: 1,141 · Forks: 111
- Language: Python
- License: MIT
- Published: 2026-09-30 · Updated: 2026-09-30 · Language: en
- Canonical page: https://hysenlabs.com/projects/awlevin-typesafe-computer-use

## The cost gap is output pricing, because the inputs are the same size

The comparison table is the centre of this project's argument, and it has one line that changes how you read the rest. Input tokens are listed as 4,882 for this loop and 4,785 for a frontier model on a bare screenshot, which is the same, and the input-token row is labelled as much. Everything else follows from output: a decision costs two ten thousandths of a dollar against three and a half cents, latency is a fraction of a second against five, and the multipliers run from 155x on a single decision to three hundred and ninety times in a loop that carries history. The mechanism is named elsewhere on the page: the classifier returns free output tokens while the frontier model writes a plan. So this is not a claim that the small model reads a screen more efficiently; it is a claim that a decision does not require prose. The scope of the measurement is equally narrow, one decision on the same screenshot and goal, which is not the same as a task that finishes.

## What the big model does for free, you rebuild as code

The page states its own central cost in a paragraph most projects would leave out. In the demonstration, the frontier model read event dates off the pixels and compared them unaided. The classifier could not, and needed the date parsing described in the step documentation to do it. Every piece of reasoning the large model performs implicitly has to be rebuilt as deterministic state, and that is the trade the whole design rests on. Step two of the loop is where the rebuilding shows up: code adds the date on any block and how far off it is, the row of a repeated label, which field is focused, the app and URL, and which actions have already been tried on this screen. The last of those is what stops the loop repeating itself, and it is a few lines of bookkeeping standing in for judgement a large model would supply without being asked.

## Your terminal is on screen, so its text is OCR input

The first run instructions contain one sentence that tells you more about the security surface than any warning does: clear the terminal first, because it is on screen and its text is OCR input. The five documented invocations differ only in flags:

```bash
uv run clicker "open the Playground"                 # dry run: one step, prints what it would do
uv run clicker "open the Playground" --act           # drives the machine, up to 100 steps
uv run clicker "log in" --act --steps 20 --delay 3   # longer and slower
uv run clicker "log in" --act --handoffs 0           # the classifier alone: its first stop ends the run
```

So the step count, the delay between steps, and whether the writer is ever consulted are all command line switches, and the default is the safe one: without the act flag the tool prints what it would do and stops. The perception layer is deterministic, a Vision OCR pass over a crop of the frontmost window plus the labelled controls from the accessibility tree, which is there because it sees icons that OCR cannot. So every scrap of text on that display becomes model input, including whatever another application has drawn there. What bounds the damage is the shape of the decision: one request answers three choices, which kind of action, which item, which site, and the action has to come from that space. The page is explicit that the classifier picks and the writer never picks an action; when the classifier stops, the writer reads the screen and either answers, hands the run back with a single move, or asks you a question. So the injection surface is real and the response surface is enumerated.

## One permission fails silently, and the panic control is a mouse slam

The macOS setup asks for two permissions for the terminal, and describes the failure mode of each in one sentence. Without screen recording, captures are wallpaper, which at least looks obviously wrong. Without accessibility, synthetic clicks are silently dropped, and the flag that actually drives the machine refuses to start, so the tool catches the dangerous half of that failure. Stopping a live run takes two forms, and one of them is worth knowing about before you need it: interrupt it in the terminal, or slam the mouse into the top left corner of the screen. The loop also ends itself when it decides it is done, when confidence drops, when it stalls, or when it reaches the step limit, and the writer then takes over to read the screen and print an answer. Every run writes a timestamped folder, which is what makes a stall replayable offline rather than merely lost.

## Almost every dependency is frozen exactly, including two zero version Windows packages

The dependency list uses exact pins for almost everything. The two model clients are pinned, the classifier SDK is pinned, the websocket client is pinned, and the platform packages are pinned per operating system: a Vision OCR binding and two Objective-C bridges on macOS, and a process library, a Windows automation library, and an OCR package on Windows. Pillow is the only entry with a range. The install is a clone, a dependency sync, and a copy of the example environment file, and that example file is where the writer model choices are visible: a small model named as the default for writing, a larger one named for answering, and switches to send the writer to another endpoint speaking either of the two common message formats, with a comment promising that only one key is ever sent there. That is a defensible choice for a tool whose entire premise is reproducible step costs, and it has a cost too, because a security fix inside a pinned client is a change to this file and a new release rather than an upgrade picked up automatically. The Windows column is where the pins look least settled: the OCR package there is at a zero version and the automation library is too, on a platform the page itself calls experimental, while the macOS pins are at least on version one.

## The benchmark runs on Linux, and the product is macOS first

There is a benchmark integration with a desktop task suite, where this loop runs as one agent beside the suite's own agent on the same task. Two commands matter: one fetches the suite and installs this loop beside it, and the other runs a single task against Chrome with a named OCR backend. Both need a Linux host with hardware virtualisation, and the documentation points at a one command machine in a cloud provider for anyone without one. The dependency list explains why, in a comment attached to an optional extra holding a portable OCR engine and its runtime: the macOS Vision framework does not exist on Linux, the benchmark reads the virtual machine's screen with the portable engine instead, and the extra is optional precisely so a Mac on Vision never installs it. The consequence is that the benchmarked configuration is the portable one. What gets measured is not what most users will run.

## The sandbox fixes one project name and binds its screen to loopback

There is a compose file for running the agent against a Linux computer in a container, for the cases where the screen it drives is not yours. Three decisions in it are worth reading. The port is published on the loopback address only, with a comment saying the graphical session is reachable from this machine only. The environment file holding your keys is passed in but marked optional, so the container starts without it. And the project name is fixed rather than derived from the directory, with the comment explaining that every checkout and worktree should share one sandbox on that port. That last one is a collision by design: two worktrees cannot each have a sandbox, and the trade is that you never accumulate stray machines. Two smaller touches are sensible, a memory backed temporary filesystem so browser profiles and scratch files never reach a disk, and an enlarged shared memory segment because the browser draws through it.

## Two benchmark directories, two assistant instruction files, and a 0.2.0 beta

The root inventory explains some of the project's habits. There are two directories whose names differ by one letter, one singular and one plural, both plausibly holding benchmarks. There are two assistant instruction files at the top level rather than in a documentation directory. There is a vision statement the page points to and an observations file beside it, and the page's own framing is that the vision file says in a dozen lines what the tool is for. The package metadata marks this as a beta release with an environment classifier for both macOS and Windows, at version 0.2.0, which matches the single tag published in late September 2026 followed by a push days later. The readme asks you to expect settings and behaviour to change between pre one releases, and to start with a dry run because this drives the real mouse and keyboard. Both requests are consistent with a version number this young.

## Conclusion

Adopt this if your automation is a sequence of choices from a known list, where a fifth of a cent a step changes the economics, and start with the dry run because it drives a real mouse. Two things to weigh before you trust it. Your screen is the input, so anything visible to the camera layer, including your own terminal, becomes text the system reads; and every reasoning step a frontier model does implicitly has to be rebuilt here as code you maintain. Also check the pins, because nearly every dependency is frozen at an exact version and the experimental Windows path carries a zero version OCR package.

## FAQ

### What is typesafe-computer-use?

It is a Mac automation loop, nicknamed jev, that reads the screen deterministically with the operating system's OCR and accessibility tree, asks a small decision model which action comes next from a short list, and calls a writing model only when free text is needed or the classifier has stopped. The rule it runs on is that the classifier picks, code decides facts, and the writer only writes text.

### How much does typesafe-computer-use cost per step?

The page puts one decision at two ten thousandths of a dollar against three and a half cents for a frontier model on the same screenshot and goal, about 155 times cheaper, with input tokens nearly identical at 4,882 against 4,785. The difference is that the classifier's output tokens are free while the larger model writes a plan. End to end a step is about 1.5 seconds against about 5.5.

### What permissions does typesafe-computer-use need on macOS?

Screen recording and accessibility, both granted to your terminal. Without screen recording, captures are wallpaper. Without accessibility, synthetic clicks are silently dropped, and the flag that drives the machine refuses to start.

### Does typesafe-computer-use work on Windows?

Windows 10 and 11 are described as experimental, with their own adapter and a documentation page explaining how it differs from macOS. The package carries pinned Windows dependencies, including an OCR package still at a zero version.

### Can I dry-run typesafe-computer-use before it touches my machine?

Yes. The same command without the act flag prints what it would do for one step. Every run also writes a timestamped run folder, so a stall can be replayed offline, and a separate inspect entry point gives a countdown capture plus an annotated screen and the payload.

## Sources

- [awlevin/typesafe-computer-use on GitHub](https://github.com/awlevin/typesafe-computer-use)
- [Issues](https://github.com/awlevin/typesafe-computer-use/issues)
- [License: MIT](https://github.com/awlevin/typesafe-computer-use/blob/main/LICENSE)
- [README](https://github.com/awlevin/typesafe-computer-use/blob/main/README.md)
- [Releases](https://github.com/awlevin/typesafe-computer-use/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/awlevin-typesafe-computer-use
