Layered memory instead of a longer transcript, and a licence gate on an MIT package
Autonomous self-evolving agents. Vision-grounded layered memory and self-written skills for LLM agents that operate your computer.
At a glance
- What is it?
- Photo Agents is a Python package that drives a tool-calling model through a perceive, reason, act loop with vision input and four tiers of memory on disk, self-written skills indexed by vectors, sandboxed execution in three shells and a DevTools bridge for browser control, and it runs on four runtime dependencies while requiring a remotely validated API key that its own launcher and hub will refuse to start without.
- Who is it for?
- It fits a developer experimenting with agents that act on a real desktop, since the memory tiers are plain files you can read between runs and the whole runtime is four dependencies you can audit. Four things to check first.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 91 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Memory is layered on disk instead of appended to a transcript
The design argument is made against the obvious alternative, and it is worth restating because it explains the whole file layout.
The stated alternative is dumping longer chat transcripts into a model and hoping for the best. Instead, vision goes in, observations are bound and stored in layers, and skills are written by the agent itself from real success. The comparison drawn is with how biology handles memory, which is a claim about consolidation rather than about storage size.
The layers are named in two places, and putting them together gives the scheme. The feature list describes a layered memory system with four roles: working, global, standard operating procedure, and session archive. The on-disk table then shows where two of them live. Long-term facts are a single text file for the global layer, described as L2. Raw session archives are a directory of files, described as L4. The skills directory holds L3 standard operating procedures and helper modules. The working layer is the one with no file of its own, which fits, since working memory is meant to be per task.
So the practical effect is that a run's raw transcript is preserved on disk as an artefact, while what carries into the next run is a short facts file and a set of procedures. That is a different failure mode from a growing context window: when the agent behaves oddly you can open the facts file and see what it believes.
The self-evolving part is a scheduler you supply yourself
Everything labelled self-evolving routes through one directory and one command, and the command takes a path from you.
The reflect mode is invoked by pointing the runtime at a script inside the evolution directory. The documentation describes it as reflect or watchdog mode, where a check function you write fires the next task. So the loop is not a built-in scheduler with configuration; it is a hook, and the thing that decides what runs next is a function in a file you point at.
That directory is described in the layout as reflection and scheduler scripts, and it is labelled as the self-evolving loop. Adjacent to it is a skills directory holding the standard operating procedures and helper modules for browser and vision work, plus a vector index directory used to search skills and procedures.
The search mechanism is the interesting part. Skill retrieval is a vector index on disk rather than a prompt dump, which is what makes the procedure set grow without the context growing. So the evolution story is: an agent writes a procedure, the procedure is indexed, and a later task retrieves it instead of re-deriving the approach.
There is a separate scheduler described as cron-style in the feature list, sitting alongside the optional tracing integration. So there are two scheduling surfaces, the reflection hook you supply and a recurring scheduler, and the documentation does not say how they relate.
Four runtime dependencies, and no vendor SDK anywhere
The dependency list for a package that controls a computer, drives a browser and routes between model providers is four lines long.
An HTTP client, an HTML parser, an XML parser, and an image library. All four are floored rather than pinned. That is the whole runtime requirement list.
The absence that matters is the missing model libraries. There is no vendor SDK in the dependency tree, which means the multi-provider router speaks to providers over plain HTTP itself rather than delegating to a maintained client for each one.
The consequences cut both ways. On the good side, adding a provider does not mean adding a dependency, and the package installs on any machine with a network client. On the other, protocol handling, retry behaviour, streaming formats and error shapes are the project's problem, and they will not be fixed upstream for you when a provider changes its API.
The router's own structure is described in the README and in the layout: native support for two major providers, plus a mixin session that fails over between them. And the troubleshooting section reveals how the selection actually works, which is by keyword rules in the credentials file rather than by anything clever: a keyword pairing of native with a provider name, or mixin.
So provider choice is a string match in a config file you edit. If two provider names both appear in your config, the wrong one can be picked, and the documented fix is to check the keyword rules.
An MIT package that will not start without a remote key
There is a licence gate, and it is the first thing that differs from what the licence text implies.
The whole thing is gated by a remotely validated API key, with the stated purpose being that usage stays accountable. Every entry point checks it. The launcher and the service hub both call the same gate before starting anything and will refuse to launch if the key is missing or revoked.
The key can arrive three ways, and the order is fixed. An environment variable first, then a saved configuration file under the user's home directory, then an interactive prompt on first run which offers to save the value. A successful validation is cached for twenty-four hours, with the stated reason being that the gate stays fast.
So the shape is: MIT source, published on a package index, installable with one command, and gated at runtime by a remote check. Those are not contradictory, but they are a choice, and the honest reading is that the licence covers the code while the service terms cover access.
The privacy answer in the FAQ is consistent with that design and worth quoting accurately. Asked whether screen data goes anywhere besides the chosen provider, the answer is no: the runtime talks to the provider and to the licence endpoint, and nothing else. So the key check is a licence check, not a data path.
The other answer in that section is narrower than you might hope. Fully offline use is not available, because the agent loop itself needs a network-reachable provider, even though memory and skills are stored locally.
Six frontends, each launched as its own module
The clients are optional extras, and the launch table is the whole integration surface.
Installation is one command, or one command with every optional client and integration included:
pip install photoagents
# or, with every optional client and integration
pip install "photoagents[all]"A Streamlit web app with an embedded webview runs through the CLI launcher module. A service hub with start and stop runs through its own module. A PyQt desktop app, a second desktop companion, a Telegram bot, and one client per Chinese workplace platform, covering Feishu, WeCom, DingTalk and QQ, each launch as their own module.
Two details in that table are worth noting. Some entries are launched with a windowless Python interpreter and some with the normal one, which on Windows is the difference between a console window appearing and not appearing. And the messaging platforms are separate optional dependencies rather than one shared abstraction, so installing the Telegram support does not give you the others.
The feature list also mentions a streaming agent loop and a physical-execution toolset, and the toolset is the part with a real risk profile. File input and output, sandboxed code execution in three shells, and browser automation through a bridge to the Chrome DevTools Protocol.
Browser control via a debugging protocol means the agent drives a real browser with a real session rather than a headless renderer, which is what makes page inspection possible. It also means the agent can see whatever is logged in in that browser profile.
The troubleshooting section confirms the bridge is a local configuration artefact rather than a managed service: if browser tools are not working, check that the parser library is installed and that the bridge configuration exists under a specific directory inside the resources folder.
State is five paths, all readable and all deletable
The on-disk table is short, and everything the agent knows lives in it.
A configuration file under the user's home directory holds the API key and the licence validation cache. A plain text file next to it holds long-term facts, the global layer. A sessions directory holds raw session archives, which is the L4 tier. A skill index directory holds the vector index used to search skills and procedures. A temp directory holds per-task scratch, described as logs and intermediate output.
Five paths, no database, no encrypted store, no hidden state.
That is a real property for an agent runtime. You can read what the agent believes, you can diff the facts file between runs to see what it learned, and you can delete one session archive without touching anything else. For a research tool that is exactly what you want, and for a production tool it is also the weakness, since there is nothing enforcing retention and nothing protecting the files from other processes on the machine.
The configuration file holds a credential in plaintext, which is worth naming explicitly. It is the standard trade for a single-user local tool, and the reason to know it is so you can set permissions on the home directory rather than assume the file is protected by the application.
The skills directory is the other one to look at, because a procedure written by an agent is code-shaped content that will be retrieved and followed later.
Beta, a root-level package marker, and an OCR extra
Three structural details tell you what stage this is at.
The version is below one and carries a comment marking it as beta in the manifest. The status section says beta and warns that interfaces may change before one. The classifiers agree, marking the development status as beta and the intended audience as developers. There are no tagged releases, so there is nothing to pin to and no upgrade notes beyond a changelog file.
Second, the repository root contains package initialisation and entry point files sitting next to the package directories, rather than inside a single package directory. That means the checkout is importable in a way that the published wheel is not necessarily structured to match, which is a common source of confusing behaviour when you run from a clone rather than from an installed package.
Third, the extras tell you what the project considers peripheral. A development extra brings a test runner, a linter and a type checker. A web extra brings an embedded webview and the dashboard framework. A desktop extra brings the Qt binding, with a Windows-only package added conditionally. A tracing extra brings the observability integration. And a vision extra is the only one that pulls real weight: an OCR runtime, an object detection library and an array library.
The all extra is the union of those minus the development tools. And it does not include the Windows-conditional package, so a Windows install of everything still needs a second step.
Editorial conclusion
It fits a developer experimenting with agents that act on a real desktop, since the memory tiers are plain files you can read between runs and the whole runtime is four dependencies you can audit. Four things to check first. It is a beta with a version below one and an explicit warning that interfaces may change. Installing it requires a key validated against a remote endpoint, and validation is cached for twenty-four hours, so the package is MIT in licence but not self-contained in operation. The router talks to providers over plain HTTP rather than through vendor libraries, so there is no SDK in the dependency list to keep current for you. And full offline use is not possible, because the loop needs a reachable model even though memory and skills never leave the machine.
Frequently asked questions
What is Photo Agents?
It is an MIT-licensed Python package providing a perceive, reason, act agent loop for any tool-calling model, with vision-grounded layered memory, self-written skills, sandboxed code execution, browser automation through a Chrome DevTools bridge, and pluggable clients. It requires Python 3.10 or newer and is documented as beta with interfaces that may change before 1.0.
How does Photo Agents store memory?
In layers on disk rather than in a growing transcript. Long-term facts sit in a single text file described as the L2 tier, raw session archives in a sessions directory described as L4, and skills and standard operating procedures in a skills directory at L3, indexed by a separate vector index directory. Working memory is per task and has no file of its own.
Does Photo Agents send my screen data anywhere?
The documented answer is no. The runtime communicates with your configured language model provider and with the project's licence endpoint, and nothing else. Fully offline operation is not possible, because the agent loop needs a network-reachable provider, even though memory and skills are stored locally.
Why does Photo Agents require an API key for an MIT-licensed package?
The whole runtime is gated by a remotely validated key, with the stated reason being that usage stays accountable, and every launcher refuses to start without one. The key comes from an environment variable, a saved config file, or an interactive first-run prompt, and a successful validation is cached for twenty-four hours so the check stays fast.
What are Photo Agents' runtime dependencies?
Four: an HTTP client, an HTML parser, an XML parser and an image library, all with lower-bound versions. There is no vendor model library in the dependency list, because the multi-provider router talks to providers over plain HTTP itself, choosing between native and failover modes through keyword rules in a credentials file you copy from a template and edit.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/jmerelnyc-photo-agents)