activity-frames: compiling local screen capture into agent-executable workflows
Turn your workday into structured workflows agents can execute. 100% local, served over MCP.
At a glance
- What is it?
- activity-frames is a Python package that records your screen locally, compiles the capture into bounded activity frames, and serves them to agents over MCP. The idea is sound; the platform story and the capture dependency are where you should look first.
- Who is it for?
- Adopt activity-frames if you are building a computer-use or desktop agent on macOS or Linux, you are comfortable with a recorder watching your screen, and you want your agent to start each session knowing what you did rather than asking. Do not adopt it if you need Windows support, if your workflows cross machines, or if you cannot accept that the quality of every compiled frame is bounded by what the recorder captured.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 35 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap activity-frames is aimed at: agents that start every session blind
A computer-use agent re-derives each task from scratch. Screenshot, reason, act, repeat. If you ran the same LinkedIn search yesterday, the agent still pays the full reasoning cost today, because nothing in its context says you have done this before. The README states the problem directly: "Computer-use agents work every task out from scratch, even one you've done a hundred times." The second half of the claim is about context. Between tasks the agent has no record of your day, so it begins every conversation with nothing.
activity-frames targets both halves with one capture pipeline. The audience is narrow and specific: developers building agents that act on a desktop, and people who want an agent to answer questions about their own working day without shipping that day to a vendor. The package is pure Python, MIT licensed, and declares no runtime dependencies in pyproject.toml. Local is the design constraint, not a feature bullet.
From snapshot rows to activity frames: the two-tier contract
The mechanism has three stages. Capture writes instants: the README describes "thousands of snapshot rows a day, each one saying 'at 22:53:05, Chrome showed linkedin.com/in/...'." Those rows are deliberately dumb. On their own they are not something a model can reason over.
The compiler turns instants into activity frames. A frame is a bounded record: an app, a site, a start and end time, a duration, a list of typed pages, input counts, and an evidence range pointing back at the snapshot rows it came from. The README's example frame carries `evidence: {frame_ids: "99871..100147"}`, which is the part that matters. Every compiled claim can be traced to the raw capture that produced it.
SPEC.md defines a two-tier contract. Tier 1 is the measured tier, and the README says everything in it is "derivable by deterministi" (the sentence is cut off in the README as published). The intent is clear from the rest of the document: the compiler does not interpret, it aggregates. The context block the README shows is labelled "measured from screen capture, no interpretation". That constraint is what makes the output usable as agent context. A frame that says you spent 18 minutes on linkedin.com doing a people search is a fact. A frame that says you were recruiting is a judgement, and the package does not make it.
Installing activity-frames and getting a first context block
Installation is a single pip command. The package declares no runtime dependencies, so there is nothing else to pull in for the basic path. PyYAML is an optional extra under the `yaml` key, which you want if you intend to parse frame output as YAML rather than read it as text.
pip install activity-framesThe `aframes` console script is declared in pyproject.toml under `[project.scripts]` and maps to `activity_frames.cli:main`. Start capture with the record command. The README notes audio is off by default, which is the right default for a tool that watches your screen.
aframes recordLet it run through some real work. Then ask for a context block covering your recent activity. The README's example uses a two-hour window.
aframes contextWhat you should see is a compact text block: a coverage line with the time range, active minutes and app count, an away period, then one line per frame with time range, app, site, duration and a short page summary. The README reports that a real day with 44 frames compiles to 1,371 tokens against 247,563 raw snapshot rows, in under a second, with no LLM in the loop. Those are the project's own numbers for one day, not a general benchmark. Verify the shape of your own output before you build on it.
The second entry point is workflow extraction. Point `aframes steps` at a task you have performed and it returns an ordered replay view.
aframes steps --find "message john doe"The README shows the result as JSON: a `steps` array where each entry has a timestamp, an operation such as `focus`, `click` or `type`, a target grounded by element name and role, a URL where relevant, and an `n` counter. The example ends with `"step_count": 6` and `"unresolved_clicks": 0`. That last field is worth watching. A nonzero value means the compiler saw a click it could not attribute to an element, and a replay built from that frame has a hole in it.
The recorder is the ceiling on everything downstream
Every frame is an aggregation of what the recorder captured. If the recorder misses a window, or captures a page state before it settles, the frame is wrong, and the frame is what your agent reads. The evidence range does not fix this. It lets you go back and check, which is a debugging affordance, not a correctness guarantee.
The platform constraint is sharper. The pyproject classifiers list MacOS and POSIX Linux. Windows is not claimed. If your team is on Windows, this is not a tool you can evaluate and adopt; it is a tool you would have to port. The README does not document a Windows path, and the repository's top level gives no separate recorder package for it.
There is also a scope boundary the README is honest about. On the happy path a compiled workflow replays at zero model calls. Anything unexpected halts and asks rather than guessing. That is the correct safety posture, and it also means the token savings apply to the routine case only. A workflow that hits an interstitial, a cookie banner or a changed layout stops and hands control back. You get reliability on repetition and no magic on novelty.
How this differs from conversation memory and from raw screen recorders
The obvious alternative is the memory layer already sitting in most agent stacks: conversation memory, which stores what you told the model. The README draws the line itself, calling what you actually did "the missing half". The difference is structural, not incremental. Conversation memory is authored by the user and shaped by the chat interface. activity-frames is authored by the machine and shaped by the screen. A question like "what was that article I had open around 9pm?" is answerable from the compiled block and unanswerable from a chat transcript.
The second alternative is a plain screen recorder. Tools in that category produce video or periodic screenshots. They preserve everything and structure nothing, which is why an agent cannot use them as context: the token cost of a day of screenshots is the token cost of the day. activity-frames trades fidelity for structure on purpose. You lose the pixels and keep the typed pages, the input counts and the time bounds. The README's own comparison makes the trade visible: 247,563 raw rows become 1,371 tokens. That is a compression ratio, and compression always discards.
The third point of comparison is the project's own research directory. The README describes a deterministic executor in `research/` that replays a compiled workflow in a real browser, plus measurements of a Routine Overhead Ratio on weeks of real activity and a public web-task dataset. pyproject.toml is explicit that the published source distribution excludes `/research`, `/paper` and `/benchmark`. The pip package is the library and docs. If your plan depends on the executor, you are reading the repository, not the release.
Maintenance, distribution and what the MIT licence leaves you to handle
The last push to the default branch was on 2026-08-26, and the repository is not archived. The most recent release listed is v0.2.2 from 2026-07-28, following v0.2.0, which the release name ties to a communications view. The package classifier marks Development Status 4 - Beta, and the version numbering is consistent with that. Treat the surface as moving.
Upgrade cost is low in the dependency sense and not zero in the data sense. There are no runtime dependencies to reconcile, and Python 3.9 through 3.13 are supported, so the package will not pin your interpreter. The thing that changes between versions is the frame schema and the compiled context format. If you persist frames or feed the context block into a stored prompt template, a schema change is your migration, and the CHANGELOG is where you would look for it. The README does not document a rollback path for the recorder or a schema migration command.
MIT is permissive. You can use it commercially, modify it and redistribute it, subject to the usual attribution requirement. That is the licence grant, not legal advice about your situation. The part worth thinking about is not the licence but the data. The capture is local, which keeps the raw material on your machine, but the compiled context block is designed to be dropped into a system prompt. Once it is in a prompt, it leaves your machine if your model is remote. The package keeps capture local; it does not control what you do with the output.
Editorial conclusion
Adopt activity-frames if you are building a computer-use or desktop agent on macOS or Linux, you are comfortable with a recorder watching your screen, and you want your agent to start each session knowing what you did rather than asking. Do not adopt it if you need Windows support, if your workflows cross machines, or if you cannot accept that the quality of every compiled frame is bounded by what the recorder captured. Before you commit, verify three things in your own environment: that aframes record produces frames with the apps you actually use, that the MCP server connects from your client, and whether your replay path needs the deterministic executor in research/, which the published sdist excludes.
Frequently asked questions
What Python version does activity-frames require?
pyproject.toml sets requires-python to >=3.9, and the classifiers list Python 3.9 through 3.13. There are no runtime dependencies, so the package does not constrain your other libraries.
Does activity-frames run on Windows?
The pyproject classifiers list MacOS and POSIX Linux only, and the README does not describe a Windows path. On Windows you would be porting the recorder rather than installing the package.
How do I start recording and get an agent-ready context block?
The README gives two commands: run aframes record to start capturing, with audio off by default, then aframes context to produce a block covering your recent activity. The README's example uses the last two hours.
What is an activity frame?
An activity frame is a bounded record of a task you performed, holding the app, site, start and end time, duration, typed pages, input counts and an evidence range pointing back to the snapshot rows it was compiled from. The README describes capture rows as instants that are useless to reason over on their own, with frames as the compiled form.
Does activity-frames send my screen activity to a server?
The capture is local, and the README states the context block compiles with no LLM in the loop and at zero token cost. The compiled block is designed to be dropped into a system prompt, so where it goes after that depends on the model you point at it.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/nossa-y-activity-frames)