Model or dataset
ClaudioDrews/memory-os avatar
ClaudioDrews/memory-os

Memory OS for Hermes Agent: a 7-layer local memory stack built on Qdrant and SQLite

A 7-layer memory operating system for Hermes Agent — persistent memory with Qdrant, structured facts, fabric recall, auto-curated wiki, and surgical context injection. Runs locally, any LLM provider.

1,373 stars128 forksPythonMIT

At a glance

What is it?
Memory OS gives Hermes Agent persistent memory across sessions through seven layers, from workspace files to a Qdrant vector store. It runs locally and works with any LLM provider, but it assumes you already run Hermes and Docker.
Who is it for?
Adopt Memory OS if you already run Hermes Agent, keep Docker on the same machine, and want memory that never leaves your hardware. Do not adopt it if you need memory for a non-Hermes agent, or if you cannot run Qdrant, Redis and the ARQ worker alongside your agent.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 10 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem Memory OS addresses for Hermes Agent users

Hermes Agent sessions are stateless by default. You configure preferences, settle on a stack, work through a hard problem, and the next session starts from zero. The README lists the symptoms directly: repeating context at the start of every conversation, losing the thread of decisions made weeks ago, and having nowhere durable for structured facts about your stack and projects to live. That is the gap Memory OS targets.

The intended user is a Hermes Agent operator running the agent on their own machine, not a team evaluating a hosted memory API. The README frames the motivation as months of hitting these limits in production, and the design follows from that: memory infrastructure runs entirely on your machine, and the LLM provider behind Hermes can be OpenRouter, OpenAI, Anthropic, Ollama or a local model. The project is explicit that there is no memory subscription and no vendor lock-in on the memory side. If you do not run Hermes Agent, this project has nothing to attach to.

How the seven layers fit together and when each one fires

The architecture is a stack of storage tiers with different retrieval characteristics, not one database with a schema. Layer 1 is flat files (MEMORY.md, USER.md, CREATIVE.md) injected into the system prompt on every turn. Layer 2 is state.db, SQLite with FTS5, giving full-text search over conversation history. Layer 3 is memory_store.db, which holds structured facts with entity resolution, HRR, FTS5 and trust scoring, plus a feedback loop that adjusts those scores over time. Layer 4 is the Icarus plugin, a fork that extracts sessions with an LLM and exposes 16 tools including fabric_recall, fabric_write and fabric_brief. Layer 5 is Qdrant with 4096-dimensional cosine vectors plus BM25 sparse vectors, and it is where the fallback logic lives: hybrid search, then dense, then lexical, then SQLite. Layer 6 is an auto-curated wiki (concepts/, entities/, comparisons/) continuously ingested into Qdrant. Layer 7 is the identity layer: SOUL.md and rulebook.md.

The data flow is two hooks. On pre_llm_call the system performs surgical recall from four sources (Fabric, Qdrant, Sessions, Facts) and injects the result. On post_llm_call and on_session_end it extracts and captures new learning. Each source is gated by relevance thresholds, per-session deduplication stops the same context appearing twice, and a social-closer filter skips trivial messages. That gating is the part worth judging: the quality of your memory depends on thresholds you will have to tune, and the README does not publish recommended values.

Layer 7 is the project's own claim to novelty, and the argument is specific. Without it, the README states, Qdrant points get injected but the agent calls the Qdrant API to verify them, Fabric entries get injected but the agent runs fabric_recall to re-find them, and facts get injected but the agent probes fact_store to confirm them. The project calls the result memory-zero behavior despite perfect injection. Whether you accept the framing or not, the underlying observation is real for any RAG-style injection: retrieved context is not automatically trusted context.

Installing Memory OS and getting a first session to remember

The requirements are Hermes Agent, Docker (Qdrant, Redis and an ARQ worker), and Python 3.11 or newer. The v0.2.0 release notes describe a one-command install that sets up Docker services, SQLite databases, the Icarus plugin and the environment. The README gives the command as:

bash
curl -sSL https://raw.githubusercontent.com/ClaudioDrews/memory-os/main/setup.sh | bash

Piping a remote script into bash is the documented path, so read setup.sh before running it if that matters to you. The release notes say a 10-step manual guide remains as a fallback for troubleshooting, which is the route to take if the automated installer fails.

Configuration is environment-driven. Copy the example file and fill it in:

bash
cp .env.example .env

Several values must be absolute paths, and the example file says why: systemd does not expand ~. The keys you will set include REDIS_PASSWORD (the file suggests generating it with openssl rand -hex 16), FABRIC_DIR, VAULT_PATH, WIKI_ROOT, HERMES_HOME, STATE_DB_PATH and HERMES_LOGS_DIR.

Embedding is a separate choice. The example file presents Ollama as the local, zero-cost option and OpenRouter as the cloud alternative, and notes that if both are set, Ollama takes priority:

bash
OLLAMA_BASE_URL=http://localhost:11434
OLLAMA_EMBEDDING_MODEL=nomic-embed-text

Host-side Python dependencies are listed in requirements.txt, which pins qdrant-client>=1.17.0, fastembed>=0.4.0, arq>=0.28.0, redis>=5.0.0, pyyaml>=6.0 and python-dotenv>=1.0.0 among others. Install them with:

bash
pip install -r requirements.txt

For a first real use, the check is behavioural rather than a command: hold a session with Hermes in which you state a preference or a decision, end it, then start a new session and see whether the agent already knows it. If it does not, the README's own diagnostic applies. Memory may be captured and injected while the agent still ignores it, which points at Layer 7 rather than at Qdrant.

Where Memory OS breaks down or is the wrong choice

The dependency surface is the first constraint. Qdrant, Redis and an ARQ worker all run under Docker, so this is not a plugin you drop into a Python environment; it is a local service stack. On a modest machine, the release notes acknowledge that Docker build times exposed UX gaps during testing, which were then handled gracefully. Graceful handling is not the same as a fast install.

The second constraint is provider scope. Memory OS is built for Hermes Agent. The Icarus plugin, the pre_llm_call and post_llm_call hooks, and the Ground Truth files all assume Hermes's extension points. If you run a different agent framework, none of the seven layers have anywhere to hook in.

The third is operational complexity that the README does not resolve. There are dead-letter-queue paths (HERMES_DLQ_PATH, HERMES_DLQ_REPORT_LOG), a telemetry log, and a reflection trigger log, which tells you the maintainers expect ingestion failures and slow reflection cycles in practice. The README does not document rollback, nor does it state what happens to existing Qdrant collections when you change embedding models. Switching from nomic-embed-text to a different model is not a documented migration path.

Finally, the relevance thresholds that gate each recall source are described as existing, not as tuned defaults. A system with four retrieval sources and per-source gates has more knobs than a single vector store, and the README does not say how to set them.

Memory OS compared with plain RAG over a vector store

The nearest alternative approach is the one most Hermes users already have: put documents in a vector database, embed them, and retrieve by similarity at query time. That is essentially Layer 5 of Memory OS on its own. The difference is what surrounds it.

A plain vector store has one retrieval mode and one relevance notion. Memory OS adds a fallback chain inside Layer 5 (hybrid, dense, lexical, then SQLite), which means a query still returns something when the embedding backend is unreachable or when the exact term matters more than the semantic neighbourhood. It adds Layer 3, where facts carry trust scores adjusted by a feedback loop rather than being ranked purely by cosine distance. It adds Layer 2, FTS5 full-text search over sessions, which is the right tool for "what did we decide about X" in a way vector search often is not. And it adds Layer 7, which is an instruction to the agent rather than a retrieval technique.

The trade-off is honest: you are running four retrieval systems and merging their output under thresholds, instead of one. If your memory needs are small and your agent already behaves well with a single vector collection, the extra layers are maintenance you will pay for without a matching benefit. If your agent keeps rediscovering things it was already told, the extra layers are the point.

Maintenance, licence and upgrade cost

The repository is not archived, and the last push was on 2026-06-10. That is roughly three months before today, so the project is not abandoned, but it is also not receiving daily commits. There are no retrieved releases beyond the v0.2.0 notes in the README, so versioning is documented through the README rather than through a release feed.

The v0.2.0 notes describe 20+ fixes from a systematic audit covering setup, configuration, performance and resilience, including provider-agnostic LLM extraction, O(1) path lookups, FTS5-powered session search, semantic dedup at scale, and idempotent database initialization. Idempotent initialization matters for upgrades: it means re-running setup should not corrupt existing databases. The notes also mention community infrastructure (issue templates, a PR checklist, a contributing guide) and state that the project already has external contributors.

Upgrade cost is concentrated in two places. First, the Docker services (Qdrant, Redis, ARQ worker) version independently of the Python host scripts, so a Qdrant upgrade is a separate operation from a Memory OS update. Second, embedding model changes are not documented as migrations; the semantic dedup threshold (cosine >0.92, per the README) and the 4096-dimensional cosine configuration in Layer 5 both assume a consistent embedding space. Changing models means re-embedding.

The licence is MIT, which permits commercial use, modification and redistribution with the licence and copyright notice retained. That is a permissive licence, but it says nothing about the licences of the forked Icarus plugin or the bundled templates and skills directories; check those separately, and treat this as a pointer rather than legal advice.

Editorial conclusion

Adopt Memory OS if you already run Hermes Agent, keep Docker on the same machine, and want memory that never leaves your hardware. Do not adopt it if you need memory for a non-Hermes agent, or if you cannot run Qdrant, Redis and the ARQ worker alongside your agent. Before committing, verify that your embedding backend is reachable (Ollama on port 11434 or an OpenRouter key), that FABRIC_DIR and VAULT_PATH are absolute paths, and that the setup script completes on your machine rather than falling back to the manual guide.

Frequently asked questions

What is Memory OS for Hermes Agent?

It is a seven-layer memory system that gives Hermes Agent persistent memory across sessions, combining workspace files, SQLite databases, a forked Icarus plugin, Qdrant vectors, an auto-curated wiki and a Ground Truth hierarchy. The README frames it as a memory operating system rather than a single plugin.

How do I use Memory OS with Hermes Agent?

The README documents a one-command install via setup.sh, which configures Docker services, SQLite databases, the Icarus plugin and the environment. You then copy .env.example to .env, set paths such as FABRIC_DIR and VAULT_PATH as absolute paths, and choose an embedding backend (Ollama or OpenRouter).

Can you use Memory OS for free?

The project is MIT licensed and the memory infrastructure runs entirely on your machine, with no memory subscription. Cost depends on your LLM and embedding provider: Ollama is described in .env.example as the local, zero-cost option, while OpenRouter is pay-per-use.

Official sources

  1. ClaudioDrews/memory-os on GitHub
  2. Issues
  3. License: MIT
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/claudiodrews-memory-os.svg)](https://hysenlabs.com/projects/claudiodrews-memory-os)