# activity-frames: compiling local screen capture into agent-executable workflows

> activity-frames is a Python package that records your screen locally, compiles the capture into bounded activity frames, and serves them to agents over MCP. The idea is sound; the platform story and the capture dependency are where you should look first.

**nossa-y/activity-frames** — Turn your workday into structured workflows agents can execute. 100% local, served over MCP.

- Repository: https://github.com/nossa-y/activity-frames
- Stars: 531 · Forks: 37
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/nossa-y-activity-frames

## The gap activity-frames is aimed at: agents that start every session blind

A computer-use agent re-derives each task from scratch. Screenshot, reason, act, repeat. If you ran the same LinkedIn search yesterday, the agent still pays the full reasoning cost today, because nothing in its context says you have done this before. The README states the problem directly: "Computer-use agents work every task out from scratch, even one you've done a hundred times." The second half of the claim is about context. Between tasks the agent has no record of your day, so it begins every conversation with nothing.

activity-frames targets both halves with one capture pipeline. The audience is narrow and specific: developers building agents that act on a desktop, and people who want an agent to answer questions about their own working day without shipping that day to a vendor. The package is pure Python, MIT licensed, and declares no runtime dependencies in pyproject.toml. Local is the design constraint, not a feature bullet.

## From snapshot rows to activity frames: the two-tier contract

The mechanism has three stages. Capture writes instants: the README describes "thousands of snapshot rows a day, each one saying 'at 22:53:05, Chrome showed linkedin.com/in/...'." Those rows are deliberately dumb. On their own they are not something a model can reason over.

The compiler turns instants into activity frames. A frame is a bounded record: an app, a site, a start and end time, a duration, a list of typed pages, input counts, and an evidence range pointing back at the snapshot rows it came from. The README's example frame carries `evidence: {frame_ids: "99871..100147"}`, which is the part that matters. Every compiled claim can be traced to the raw capture that produced it.

SPEC.md defines a two-tier contract. Tier 1 is the measured tier, and the README says everything in it is "derivable by deterministi" (the sentence is cut off in the README as published). The intent is clear from the rest of the document: the compiler does not interpret, it aggregates. The context block the README shows is labelled "measured from screen capture, no interpretation". That constraint is what makes the output usable as agent context. A frame that says you spent 18 minutes on linkedin.com doing a people search is a fact. A frame that says you were recruiting is a judgement, and the package does not make it.

## Installing activity-frames and getting a first context block

Installation is a single pip command. The package declares no runtime dependencies, so there is nothing else to pull in for the basic path. PyYAML is an optional extra under the `yaml` key, which you want if you intend to parse frame output as YAML rather than read it as text.

```bash
pip install activity-frames
```

The `aframes` console script is declared in pyproject.toml under `[project.scripts]` and maps to `activity_frames.cli:main`. Start capture with the record command. The README notes audio is off by default, which is the right default for a tool that watches your screen.

```bash
aframes record
```

Let it run through some real work. Then ask for a context block covering your recent activity. The README's example uses a two-hour window.

```bash
aframes context
```

What you should see is a compact text block: a coverage line with the time range, active minutes and app count, an away period, then one line per frame with time range, app, site, duration and a short page summary. The README reports that a real day with 44 frames compiles to 1,371 tokens against 247,563 raw snapshot rows, in under a second, with no LLM in the loop. Those are the project's own numbers for one day, not a general benchmark. Verify the shape of your own output before you build on it.

The second entry point is workflow extraction. Point `aframes steps` at a task you have performed and it returns an ordered replay view.

```bash
aframes steps --find "message john doe"
```

The README shows the result as JSON: a `steps` array where each entry has a timestamp, an operation such as `focus`, `click` or `type`, a target grounded by element name and role, a URL where relevant, and an `n` counter. The example ends with `"step_count": 6` and `"unresolved_clicks": 0`. That last field is worth watching. A nonzero value means the compiler saw a click it could not attribute to an element, and a replay built from that frame has a hole in it.

## The recorder is the ceiling on everything downstream

Every frame is an aggregation of what the recorder captured. If the recorder misses a window, or captures a page state before it settles, the frame is wrong, and the frame is what your agent reads. The evidence range does not fix this. It lets you go back and check, which is a debugging affordance, not a correctness guarantee.

The platform constraint is sharper. The pyproject classifiers list MacOS and POSIX Linux. Windows is not claimed. If your team is on Windows, this is not a tool you can evaluate and adopt; it is a tool you would have to port. The README does not document a Windows path, and the repository's top level gives no separate recorder package for it.

There is also a scope boundary the README is honest about. On the happy path a compiled workflow replays at zero model calls. Anything unexpected halts and asks rather than guessing. That is the correct safety posture, and it also means the token savings apply to the routine case only. A workflow that hits an interstitial, a cookie banner or a changed layout stops and hands control back. You get reliability on repetition and no magic on novelty.

## How this differs from conversation memory and from raw screen recorders

The obvious alternative is the memory layer already sitting in most agent stacks: conversation memory, which stores what you told the model. The README draws the line itself, calling what you actually did "the missing half". The difference is structural, not incremental. Conversation memory is authored by the user and shaped by the chat interface. activity-frames is authored by the machine and shaped by the screen. A question like "what was that article I had open around 9pm?" is answerable from the compiled block and unanswerable from a chat transcript.

The second alternative is a plain screen recorder. Tools in that category produce video or periodic screenshots. They preserve everything and structure nothing, which is why an agent cannot use them as context: the token cost of a day of screenshots is the token cost of the day. activity-frames trades fidelity for structure on purpose. You lose the pixels and keep the typed pages, the input counts and the time bounds. The README's own comparison makes the trade visible: 247,563 raw rows become 1,371 tokens. That is a compression ratio, and compression always discards.

The third point of comparison is the project's own research directory. The README describes a deterministic executor in `research/` that replays a compiled workflow in a real browser, plus measurements of a Routine Overhead Ratio on weeks of real activity and a public web-task dataset. pyproject.toml is explicit that the published source distribution excludes `/research`, `/paper` and `/benchmark`. The pip package is the library and docs. If your plan depends on the executor, you are reading the repository, not the release.

## Maintenance, distribution and what the MIT licence leaves you to handle

The last push to the default branch was on 2026-08-26, and the repository is not archived. The most recent release listed is v0.2.2 from 2026-07-28, following v0.2.0, which the release name ties to a communications view. The package classifier marks Development Status 4 - Beta, and the version numbering is consistent with that. Treat the surface as moving.

Upgrade cost is low in the dependency sense and not zero in the data sense. There are no runtime dependencies to reconcile, and Python 3.9 through 3.13 are supported, so the package will not pin your interpreter. The thing that changes between versions is the frame schema and the compiled context format. If you persist frames or feed the context block into a stored prompt template, a schema change is your migration, and the CHANGELOG is where you would look for it. The README does not document a rollback path for the recorder or a schema migration command.

MIT is permissive. You can use it commercially, modify it and redistribute it, subject to the usual attribution requirement. That is the licence grant, not legal advice about your situation. The part worth thinking about is not the licence but the data. The capture is local, which keeps the raw material on your machine, but the compiled context block is designed to be dropped into a system prompt. Once it is in a prompt, it leaves your machine if your model is remote. The package keeps capture local; it does not control what you do with the output.

## Conclusion

Adopt activity-frames if you are building a computer-use or desktop agent on macOS or Linux, you are comfortable with a recorder watching your screen, and you want your agent to start each session knowing what you did rather than asking. Do not adopt it if you need Windows support, if your workflows cross machines, or if you cannot accept that the quality of every compiled frame is bounded by what the recorder captured. Before you commit, verify three things in your own environment: that aframes record produces frames with the apps you actually use, that the MCP server connects from your client, and whether your replay path needs the deterministic executor in research/, which the published sdist excludes.

## FAQ

### What Python version does activity-frames require?

pyproject.toml sets requires-python to >=3.9, and the classifiers list Python 3.9 through 3.13. There are no runtime dependencies, so the package does not constrain your other libraries.

### Does activity-frames run on Windows?

The pyproject classifiers list MacOS and POSIX Linux only, and the README does not describe a Windows path. On Windows you would be porting the recorder rather than installing the package.

### How do I start recording and get an agent-ready context block?

The README gives two commands: run aframes record to start capturing, with audio off by default, then aframes context to produce a block covering your recent activity. The README's example uses the last two hours.

### What is an activity frame?

An activity frame is a bounded record of a task you performed, holding the app, site, start and end time, duration, typed pages, input counts and an evidence range pointing back to the snapshot rows it was compiled from. The README describes capture rows as instants that are useless to reason over on their own, with frames as the compiled form.

### Does activity-frames send my screen activity to a server?

The capture is local, and the README states the context block compiles with no LLM in the loop and at zero token cost. The compiled block is designed to be dropped into a system prompt, so where it goes after that depends on the model you point at it.

## Sources

- [Issues](https://github.com/nossa-y/activity-frames/issues)
- [License: MIT](https://github.com/nossa-y/activity-frames/blob/main/LICENSE)
- [nossa-y/activity-frames on GitHub](https://github.com/nossa-y/activity-frames)
- [README](https://github.com/nossa-y/activity-frames/blob/main/README.md)
- [Releases](https://github.com/nossa-y/activity-frames/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/nossa-y-activity-frames
