# HKUDS/CatchMe: A Local Screen Recorder That Turns Your Day Into a Queryable Memory Tree

> CatchMe records windows, keystrokes, clipboard and files, then organizes them into a five-tier Activity Tree that an LLM walks top-down instead of querying a vector database. Here is how the pieces fit, where the design strains, and what to check before running it.

**HKUDS/CatchMe** — "CatchMe: Make Your AI Agents Truly Personal"

- Repository: https://github.com/HKUDS/CatchMe
- Website: https://hkuds.github.io/CatchMe/
- Stars: 509 · Forks: 82
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/hkuds-catchme

## The gap CatchMe targets: agents that forget everything between sessions

Most agent tooling assumes the model's context is the whole world. You paste a file, ask a question, close the tab, and the next session starts from zero. CatchMe takes the opposite position: the useful context is already on your machine, in the windows you had open, the files you touched, the text you copied, and the pages you read. The project describes itself as an always-on personal digital footprint recorder with hierarchical memory and LLM-powered retrieval, and the README frames the audience as people who want their CLI agents to become personal rather than general. The four scenarios it names are coding sessions (what was I editing in Claude Code today), research (what was I reading about AI yesterday), file management (which files changed today), and a general digital life overview. The common thread is recall of your own activity, not retrieval over a document corpus you assembled deliberately. That distinction matters, because it changes what the storage layer has to do. A document corpus is written once and queried many times; a footprint is written constantly and queried rarely, in bursts. CatchMe is built for the second shape.

## Record, index, retrieve: the three stages and what each one stores

The architecture is described as three concurrent stages. Capture runs six background recorders: window focus, keystrokes, mouse movement, screenshots, clipboard and notifications. The README is specific about the trigger for screenshots: recording is event-driven rather than timer-based, and the project says mouse actions are caught instantly with a crosshair annotation rather than sampled on an interval. Around each mouse action the five context recorders collect what else was happening. Index takes that raw stream and builds a Hierarchical Activity Tree with five levels: Day, Session, App, Location, Action. Each node receives an LLM-generated summary, so the tree is not just a folder structure but a set of compressed descriptions at every level. Retrieve is where the design diverges from the usual stack. There are no embeddings and no vector database. The README states the system uses tree-based reasoning for navigation: the LLM reads the summaries, picks the branches that look relevant, and drills down until it reaches raw evidence such as a screenshot or a keystroke sequence. Storage is SQLite with FTS5, and the README puts runtime memory at roughly 0.2GB. The trade-off is visible in that choice. You avoid embedding cost, index rebuilds and the failure mode where a semantically similar but wrong chunk ranks first. You also lose fuzzy recall: if the summary of a node does not mention the thing you are asking about, the traversal has no reason to visit it, and the quality of every answer depends on the quality of summaries that were written before you knew what you would ask.

## Installing CatchMe and running a first query

The package metadata in pyproject.toml declares the distribution name catchme, requires Python 3.11 or newer, and exposes a single console entry point, catchme, mapped to catchme.run:main. That entry point is the command you run after installing. The README does not spell out a pip line in the section retrieved here, so the install path to check first is the project's own get-started section and the docs/ directory in the repository.

```bash
catchme
```

The pyproject file lists the runtime dependencies you should expect to be pulled in: flask for the web viewer, pyyaml for configuration, openai for the LLM client, mss and Pillow and numpy and imagehash for screen capture and deduplication, pynput for input recorders, psutil for system information, and pymupdf, websockets, trafilatura and requests for the document and web pipelines. On macOS the install additionally pulls pyobjc-framework-Cocoa, Quartz, ScreenCaptureKit and CoreMedia; on Windows it pulls pywin32 and comtypes. Those platform markers mean the install surface differs by operating system, and the macOS frameworks are the ones that gate screen capture permissions.

The second piece is the model. CatchMe needs an LLM to write node summaries and to traverse the tree, and the README points at a dedicated LLM configuration section. The configuration format is YAML, since pyyaml is a dependency, and the README states that offline operation is possible through Ollama, vLLM or LM Studio, with the openai client used as the interface. The exact key names are not reproduced in the section retrieved here, so take them from the README's LLM configuration section rather than guessing. After the recorder is running, the third piece is agent integration. CatchMe ships as an agent-compatible skill for CLI agents, and the README describes it as a one-file setup: you drop a single skill file into the agent, and the agent then queries memories through CLI commands only. That is the intended division of labour. CatchMe runs independently as a process; the agent never touches the database directly, it shells out.

The optional extras are declared in the same file and are worth knowing before you commit to a workflow.

```toml
[project.optional-dependencies]
dev = [
    "pytest>=7.0",
    "ruff>=0.4",
]
mcp = [
    "mcp[cli]>=1.0",
]
```

The mcp extra is the Model Context Protocol path, which is the other route an agent can take to the recorded data besides the CLI skill file.

## The privacy claim and the honest limits of a local recorder

CatchMe's pitch is that all data stays on your machine, and the README pairs that with offline mode through Ollama, vLLM or LM Studio. That combination is the strongest part of the design: if both the store and the model are local, the footprint never crosses the network. The limits are equally clear, and they are not about the software's intent. A recorder that captures keystrokes and clipboard content captures passwords, tokens, personal messages and anything else typed while it runs. The README does not document a redaction layer, an exclusion list for applications, or a pause control, and it does not document rollback or deletion semantics for captured events. The web interface is described as offering real-time system monitoring, which suggests visibility into what is running, but visibility is not the same as selective capture. There is also a legal dimension the licence does not address: recording another person's screen, or a shared machine, is governed by workplace and jurisdictional rules that Apache-2.0 says nothing about. On the retrieval side, the failure mode is quieter. Because there is no vector index, a question phrased in vocabulary that never appeared in any summary will not route to the right branch. The tree is only as good as the summaries, and the summaries are generated at capture time by whatever model you configured. A weak local model produces weak summaries, and the cost of that shows up weeks later when you ask a question the tree cannot answer.

## Where CatchMe differs from vector-based memory tools

The obvious point of comparison is a retrieval stack built on embeddings and a vector database, the pattern used by most agent memory libraries. Those systems chunk text, embed it, and at query time find the nearest vectors. CatchMe rejects that pipeline entirely: the README states there are no embeddings and no vector database, and retrieval is a top-down traversal of a summary tree. The difference in approach produces different failure modes. Vector search is good at fuzzy matching, so a paraphrase of a question usually still lands near the right chunk; it is bad at structure, because a chunk carries no information about which day, which application or which session it belongs to unless that is baked into the text. CatchMe inverts this. Structure is explicit at every level, so questions of the form what did I do in this app on this day are cheap and precise. Paraphrase recall is the weak spot. A second comparison is with screenshot-only tools that store images and run OCR at query time. Those avoid the summarization step and therefore avoid summary quality as a dependency, but they push all the reasoning cost to query time and have no notion of a session or a day. CatchMe's tree is doing that work up front. The third comparison, and the one worth weighing most, is a hosted note-taking or journaling tool. Those are cheaper to operate and require no local model, but they depend on you writing things down, which is exactly the burden CatchMe removes.

## Maintenance, packaging and licence cost

The repository is not archived. Its last push was on 2026-06-16, which is more than three months before today, so treat it as a project with a recent but not continuous commit history rather than one under constant change. Version 0.1.0 in pyproject.toml is consistent with that: this is early software, and the absence of any retrieved release notes means there is no published upgrade path to reason about. Practically, that means pinning. The dependency list includes fast-moving packages (openai, mss, pynput, trafilatura) and platform-specific native bindings, and a break in any of them breaks capture rather than retrieval, which is the harder failure to notice. The ruff configuration in pyproject.toml targets Python 3.11 with a line length of 100 and a selected rule set covering pycodestyle, pyflakes, isort, pyupgrade, bugbear and simplify, and the pytest configuration points testpaths at catchme/tests. That tells you the codebase is linted and tested to a consistent standard, not that it is stable. Licensing is Apache-2.0, which permits commercial and private use and includes an explicit patent grant; it also requires that you preserve notices and state changes. Nothing in the licence speaks to the legality of recording people, so that question is separate and depends on where you are and whose machine it is.

## Conclusion

Adopt CatchMe if you want an agent to answer questions about your own machine's activity and you are willing to run a screen and input recorder continuously on that machine, with a local model via Ollama, vLLM or LM Studio if the data must not leave it. Do not adopt it on a shared or managed machine, on a laptop you use for other people's confidential work, or if you need a hosted, multi-user memory service. Before installing, verify three things in the repository: that the recorder can be paused from the web interface, what the LLM configuration file expects, and whether the one-file skill integration is documented well enough for your agent. If the pause control is not there, treat the tool as all-or-nothing and decide accordingly.

## FAQ

### What exactly does HKUDS/CatchMe record?

The README lists six background recorders: window focus, keystrokes, mouse movement, screenshots, clipboard and notifications. Screenshot capture is described as event-driven around mouse actions rather than timer-based.

### Does CatchMe need a vector database or embeddings?

No. The README states the system skips embeddings and vector databases and instead uses tree-based reasoning, where the LLM reads summaries in the Hierarchical Activity Tree and drills down to raw evidence. Storage is SQLite with FTS5.

### Can CatchMe run fully offline?

The README says all data stays on your machine and that a full offline mode is available through Ollama, vLLM or LM Studio. The openai package is used as the client, so a local OpenAI-compatible endpoint is the expected shape.

### How do AI agents get access to the captured memories?

CatchMe runs independently and ships as an agent-compatible skill for CLI agents such as OpenClaw, NanoBot, Claude and Cursor. The README describes a one-file setup, after which agents query memories through CLI commands only.

## Sources

- [HKUDS/CatchMe on GitHub](https://github.com/HKUDS/CatchMe)
- [Issues](https://github.com/HKUDS/CatchMe/issues)
- [License: Apache-2.0](https://github.com/HKUDS/CatchMe/blob/main/LICENSE)
- [Project website](https://hkuds.github.io/CatchMe/)
- [README](https://github.com/HKUDS/CatchMe/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/hkuds-catchme
