Model or dataset
RyjoxTechnologies/Octopoda-OS avatar
RyjoxTechnologies/Octopoda-OS

Octopoda adds memory to an agent with two dependencies

The open-source memory and observability layer for AI agents — persistent memory, loop detection, hash-chained audit trails, and a live dashboard, automatic on pip install.

487 stars65 forksPythonNOASSERTION

At a glance

What is it?
A memory and observability layer that instruments an existing agent without changing its logic, installing with two Python dependencies and no services. Loop detection runs on every write while intervention stays opt-in, and the audit trail is hash-chained per agent. Its container files reference a directory the repository does not contain.
Who is it for?
Adopt Octopoda if you ship agents that forget their users between sessions and you want persistence, loop detection and an audit trail without rewriting the agent. Do not adopt it expecting a hosted service you can lean on for storage, because the local mode is a SQLite file and the cloud mode is a separate account with PostgreSQL behind it.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 83 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The core install pulls requests and pydantic

The strongest claim in the project is about what it does not need, and the dependency list is where you can check it.

The core install declares two runtime dependencies: a requests library at 2.28 or newer and pydantic at 2 or newer. Everything else is an extra. The server extra adds a web framework, an ASGI server, a second web framework with CORS, and a PostgreSQL driver. Separate extras exist for the Model Context Protocol, for LangChain, for embeddings, for a vector search library, and for a natural language processing library.

So the base install is a client, and the phrase in the README that it runs locally with one install and zero infrastructure is literally true.

The local API is deliberately small:

python
from octopoda import AgentRuntime

agent = AgentRuntime("my_chatbot")
agent.remember("user_name", "Alice")

# kill the process. restart Python. then:
print(agent.recall("user_name").value)

The example's comment is the claim under test: the value is still there after a restart. The local storage is SQLite on your machine, and the project states there is no configuration, no Docker, no Redis and no extra services.

The dashboard is one extra further. Install the server extra, run the command, and it opens on a local port, running against your local data with no account and no key.

Detection is automatic, stopping the agent is not

The loop detection design is the most interesting thing about the project, because of where it draws the line.

Detection is automatic on every write. That means you do not instrument anything to get it, and a loop that starts on the third turn is caught on the third turn.

The detector recognises five named patterns: retry, oscillation, ping-pong, reflection, and recall-write. Each is a different way an agent fails to converge, and naming them separately matters because a plain retry counter misses an oscillation where the agent alternates between two states and makes progress on neither.

Intervention is where the opt-in boundary sits. Auto-pause and a spend cap are not on by default; they are reached through the v2 circuit-breaker configuration, and the project's stated reason is that the policy stays yours.

That split is a considered choice about blast radius. An automatic pause is the right behaviour when an agent is burning money on a failing API call, and the wrong behaviour when a legitimate long-running task happens to look like a loop.

When it does fire, the output is the specific calls that caused it, not a counter. That is what turns a cost incident into something you can debug.

The audit trail is a hash chain per agent

The audit feature has a defined structure, and it is the part of this project that matters most for compliance.

Every decision, write and recovery is logged into a replayable timeline you can diff over time. Diffing matters more than logging: a log tells you what happened, a diffable timeline tells you what changed between the run that worked and the run that did not.

Events written through the audit-v2 endpoint are hash-chained per agent. Each event carries the previous event's hash and its own, so the chain is `prev_hash` to a computed `_this_hash`. Integrity is verified with a single call, which is the property that makes it useful: you are not asking a service whether it still has your logs, you are checking that the logs were not altered.

The per-agent scoping is deliberate rather than incidental. Two agents writing to one chain would let either one's activity invalidate the other's, so each agent gets its own chain.

The dashboard is where this surfaces. The live view is described as agent health, operations volume, per-agent scores, an anomaly stream, and the loops that were caught before they burned tokens, and the project states that the same dashboard runs locally and in the cloud.

Two integration paths, and one of them wraps your script

There are two ways in, and they suit different situations.

The first is a call in your process. Install the package, import it, and call the initialisation function with a key. The README describes that as the entire integration: the library auto-detects your framework, captures what matters from each turn, distills it into memories, and injects relevant recall into future calls.

The second is a wrapper that needs no code change at all:

bash
export OCTOPODA_API_KEY=sk-octopoda-...
octopoda-run python your_agent.py     # auto-instruments on launch
octopoda-run doctor                   # checks your key + detected frameworks

That is the interesting one for an existing codebase. Running an agent script through the wrapper is the whole adoption step, and the doctor subcommand checks both the key and which frameworks it found, which is the first thing to run when nothing is being captured.

Two environment variables are worth knowing. Setting an agent identifier makes several scripts write to one shared memory, which is what you want when one logical agent runs as more than one process. And a recall timeout, set to five seconds by example, exists for slow networks, which tells you recall is a network call in the cloud path and a local read in the local one.

With a key, agents and their memories are said to appear on the dashboard about ten seconds after the first turn.

The container copies a directory the tree does not contain

The deployment files describe a system that does not match the repository layout, and anyone building the container should know before they start.

The Dockerfile's first copy step targets a directory named for a Python SDK, which it installs along with the PostgreSQL driver. That directory is commented as being copied first for better caching. No such directory appears among the repository's top-level entries, which hold the package source, two other component trees, a runtime tree, a scripts directory, documentation, examples and tests.

The Compose file mounts a local nginx configuration file into the container, read-only, along with the certificate directory. That configuration file is also not in the tree.

So the container build as written would fail at the copy step on a fresh clone. That does not mean the project does not work; it means the deployment files are aimed at a different checkout or a sibling repository.

The rest of the container definition is competent and worth reading for its own sake. A health check polls the API on a thirty-second interval with a two-minute start period, which is the right shape for something that loads models on boot. The container runs as a Python slim image, installs a compiler and the PostgreSQL headers, and starts the runtime module directly with a flag to suppress opening a browser.

Two database URLs, one of which bypasses row-level security

The example environment file is more informative than it looks, and one pattern in it is worth adopting regardless of whether you use this project.

Two database connection strings are defined. One is labelled as the application role, with row-level security enforced. The other is labelled as the admin connection for migrations, with a note that it bypasses row-level security.

That is the standard separation and it is worth stating explicitly rather than leaving to inference: the connection your application runs with cannot read rows it should not, and the connection that changes the schema can. Migrations need the second kind; application code must never hold it.

The rest of the file is a list of what is optional. An email verification key, a billing key marked as future, an embedding model defaulting to a small English BGE model, and a provider and model pair for fact extraction with a local Ollama model named.

That last one is the interesting design choice. Fact extraction is the step that turns a conversation into a memory, and the default configuration points it at a local model rather than a hosted API.

The backend selector defaults to PostgreSQL and the log level to info. The file also carries the two instructions that always matter: copy it to a local environment file, and never commit that file to git.

Four component trees and an MCP directory

The repository layout suggests a project that grew past its original packaging, and that is visible in the directory names.

There is the Python package itself, a directory whose name ends in a two-letter suffix, a directory for the runtime, and one for the service. Alongside them sits a skills directory, a scripts directory, a documentation directory, an examples directory and a tests directory.

At the root there is also an MCP directory, which matches the Model Context Protocol extra in the packaging metadata, and a SQL initialisation file at the top level rather than inside a migrations directory.

The Python packaging only covers one of those trees. The package finder is configured to include the main package, so the runtime and service directories are not installed as importable packages from this project. They are reachable instead through the container, which copies two of them by name and runs the runtime module from the command line.

That is why the container's missing directory matters more than it looks. The Docker image is the supported path for the service half, and the image expects a layout this repository does not have.

The examples directory is the part that tells you what the project thinks it is for, since the files are named after situations rather than features: a local-only walkthrough, a first-five-minutes script, a loop detection demonstration, a framework comparison, a knowledge repair demonstration, a self-debugging demonstration, and two that stage a debate or a contest between models.

One release tag, and a version three minors ahead of it

The version story is worth reconciling before you pin anything.

There is a single tagged release, version 3.0.3, named as the production release, published on 2026-04-06. The packaging metadata declares version 3.3.5.

So the default branch is three minor versions ahead of anything published as a release, and the last push to it was on 2026-07-15. The README refers to a v2 circuit-breaker configuration and an audit-v2 endpoint, which suggests the versioned surfaces are still moving.

The licensing needs the same kind of check. The packaging metadata declares MIT, and the README says the whole thing is MIT-licensed, but the repository records no licence identifier in its metadata and the licence file itself is not among the entries visible here. For a project whose selling point includes a verifiable audit trail, the licence of the auditing code is worth confirming directly rather than inferring from the prose.

Beyond that, the repository carries the usual operational files: a changelog, a contributing guide, a security policy, a roadmap, and separate continuous integration workflows for tests and for a smoke run.

Editorial conclusion

Adopt Octopoda if you ship agents that forget their users between sessions and you want persistence, loop detection and an audit trail without rewriting the agent. Do not adopt it expecting a hosted service you can lean on for storage, because the local mode is a SQLite file and the cloud mode is a separate account with PostgreSQL behind it. Verify two things first. Run the local-only path and confirm the memory survives a process restart before wiring in a key. Then read the container files, since the Dockerfile copies a directory that is not in the repository tree and the Compose file mounts a configuration file that is not either.

Frequently asked questions

What is Octopoda?

A memory and observability layer that sits underneath agents written in plain Python, LangChain, CrewAI, AutoGen, the OpenAI Agents SDK or over MCP. It provides persistent memory, loop detection, a hash-chained audit trail and a live dashboard, and instruments an existing agent without changing its logic.

How do I install Octopoda?

Run pip install octopoda. The core declares only requests and pydantic as runtime dependencies. The dashboard needs the server extra, the Model Context Protocol integration needs its own extra, and a key from the project's site is required for the cloud path while local use needs no account.

Does Octopoda stop an agent automatically when it loops?

Detection is automatic on every write and covers retry, oscillation, ping-pong, reflection and recall-write patterns. Intervention is opt-in: auto-pause and a spend cap are reached through the v2 circuit-breaker configuration, so the policy stays with you.

How does the Octopoda audit trail work?

Events written through the audit-v2 endpoint are hash-chained per agent, with each event carrying the previous hash and its own, and integrity is verified with a single call. Everything else goes into a replayable timeline you can diff between runs.

Can Octopoda work without sending data to a server?

Yes. The SDK can be used directly with an agent runtime class and no account, storing memory in SQLite on your machine, with the dashboard available locally on a local port with no API key. Cloud sync and the hosted dashboard are a separate step that requires a key.

Official sources

  1. Issues
  2. Project website
  3. README
  4. Releases
  5. RyjoxTechnologies/Octopoda-OS on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/ryjoxtechnologies-octopoda-os.svg)](https://hysenlabs.com/projects/ryjoxtechnologies-octopoda-os)