Model or dataset
verygoodplugins/automem avatar
verygoodplugins/automem

AutoMem: a graph and vector memory service for AI assistants

Long-term memory for AI assistants. Graph + vector store that recalls decisions, relationships, and context across sessions.

821 stars106 forksPythonMIT

At a glance

What is it?
AutoMem stores decisions, preferences and relationships in FalkorDB and Qdrant behind a Flask API, then exposes them to Claude, Cursor, Codex and ChatGPT over MCP. Here is how the pieces fit, what the install looks like, and where the design stops being the right choice.
Who is it for?
Adopt AutoMem if you want memory you host yourself and you are willing to run FalkorDB and Qdrant alongside the Flask API; the graph is the canonical record, so a FalkorDB outage returns 503 and there is no fallback for it. Do not adopt it if a single container is a hard requirement or if you want recall to work with no embedding provider configured.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 20 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem AutoMem targets: context that dies with the chat window

Assistants forget. Each session starts from an empty context, so a decision you explained last week has to be explained again. The usual workaround is to paste history back in, which grows the prompt and does not scale. AutoMem takes the position that the useful part of that history is small and structured: a decision, the alternatives that were rejected, the principle behind the choice, and what followed from it.

The project is aimed at people who run assistants across more than one surface. The README names Claude Desktop, Cursor, Claude Code, Codex and Copilot for local use, and ChatGPT Developer Mode, Claude.ai and ElevenLabs over Remote MCP for cloud agents. If you only ever use one chat app and never revisit old decisions, the service is overhead. If you keep re-explaining the same architectural choices to different tools, the pitch is direct.

The README frames the difference from plain vector search in one line: a vector match finds something similar, while AutoMem also records typed relationships. That distinction drives everything else in the design.

FalkorDB as the source of truth and Qdrant as the index

AutoMem runs as a Flask service with two storage layers behind one API. FalkorDB holds memories as nodes connected by 11 authorable relationship types, and the README states plainly that the graph is the canonical record. Qdrant holds one embedding per memory, described as 1024-dimensional vectors.

Recall is a hybrid query rather than a single lookup. The README lists semantic similarity, graph traversal, temporal alignment, tag overlap and importance, combined into a 9-component score. Embeddings come from Voyage, OpenAI or a local provider, and the README notes that enrichment is optional and adds structure over time.

The failure behaviour is asymmetric, and this is the most useful detail in the README. If Qdrant is unavailable, the graph still serves recall in a degraded mode. If FalkorDB is down, the API returns 503. So the vector store is an accelerator and the graph is not replaceable. Anyone planning redundancy should plan it around FalkorDB first.

Multi-hop bridge discovery is the mechanism that justifies the graph. The README's example starts from two seed memories, one about migrating to PostgreSQL for operational simplicity and one about evaluating Kafka versus RabbitMQ. Both carry an EXEMPLIFIES edge to a third memory: the team prefers boring technology. AutoMem ranks that bridge above the seeds and returns it, so the assistant answers with the reasoning instead of two disconnected facts. The knobs are expand_relations, relation_limit and expansion_limit on GET /recall.

The 11 authorable types cover general connection, causality, temporal order, preference, pattern examples, contradiction, reinforcement, invalidation, evolution, source tracking and hierarchy. Three more, SIMILAR_TO, PRECEDED_BY and DISCOVERED, are written by the enrichment pipeline and the consolidation engine rather than by a person. That split matters: the edges you author are the ones you can reason about, and the ones the system adds are the ones you have to inspect when recall looks odd.

Running AutoMem locally with Docker Compose

The repository ships a docker-compose.yml with three services: falkordb, qdrant and flask-api. The FalkorDB service publishes port 6379 for the database and 3000 for its built-in graph browser, with persistence configured through REDIS_ARGS defaulting to --save 60 1 --appendonly yes --appendfsync everysec. Qdrant publishes 6333 and 6334. The Flask API maps to 8001.

Start the stack from the repository root:

bash
docker compose up -d

Once the containers report healthy, the API is reachable on port 8001. The FalkorDB browser on port 3000 is the quickest way to confirm that graph nodes are actually being written, which is worth doing before wiring up a client.

The compose file reads configuration from an env file, and the comment in the file is explicit that docker compose --env-file only affects compose-file interpolation and never the container environment. That is why the flask-api service uses env_file rather than interpolation for the settings that matter.

The Dockerfile sets QDRANT_URL, QDRANT_HOST and QDRANT_PORT, with QDRANT_PORT defaulting to 6333 and the comment noting that QDRANT_URL takes precedence when set. If you point the API at an external Qdrant rather than the compose service, set QDRANT_URL; if you are on an internal network such as Railway, set QDRANT_HOST and QDRANT_PORT instead.

bash
QDRANT_URL=""
QDRANT_HOST="qdrant.railway.internal"
QDRANT_PORT="6333"

For client-side use, the README points at the npm package @verygoodplugins/mcp-automem as the local MCP bridge. The README does not give the full bridge configuration in the excerpt available here, so check INSTALLATION.md and the package page for the exact command and arguments before assuming a shape.

Where AutoMem is the wrong tool

The clearest limitation is the dependency floor. A working AutoMem deployment needs FalkorDB, Qdrant and the Flask API, plus an embedding provider. That is three services and an external API before recall does anything useful. For a single-user setup on a laptop, or a team that wants memory in one process with no network hops, this is more moving parts than the benefit justifies.

The second limitation is the missing fallback on the graph side. The README states that a Qdrant outage degrades recall but a FalkorDB outage returns 503. There is no documented queue, no write buffer, no read-only mode for the graph. If your assistant depends on memory and FalkorDB restarts, requests fail rather than returning stale data. The README does not document rollback behaviour for a failed write, so treat writes during an outage as unverified.

The third is embedding cost and lock-in. Every memory needs an embedding from Voyage, OpenAI or a local provider. The README's claim that recall avoids an extra generative-LLM charge is about the recall path, not about ingestion; embeddings are still paid for. A local provider removes the per-call cost but adds model management to your deployment.

Finally, the benchmark numbers in the README need reading carefully. The LoCoMo figure of 85.1 percent and the BEAM 10M figure of 57.4 percent are labelled as coming from the neutral Agent Memory Benchmark, while the LongMemEval full figure of 87.0 percent is described in the badge as running on AutoMem's internal harness. Those are not the same kind of evidence, and the README itself makes the distinction. The BEAM result is the more interesting one anyway: 57.4 percent at 10 million source tokens with roughly 2.6 to 4.8k retrieved tokens given to the answerer. That is a statement about retrieval volume, not about answer quality.

AutoMem against Mem0 and plain vector search

The obvious comparison is Mem0, which appears in the related searches for this project. The architectural difference is where structure lives. A Mem0-style system typically extracts facts and stores them for similarity lookup; the graph is not the canonical record. AutoMem inverts that: FalkorDB holds typed edges as the record, and Qdrant is an index over it. The practical consequence is that AutoMem can answer why-questions by traversing EXEMPLIFIES and LEADS_TO edges, while a pure vector store returns the nearest passages and leaves the reasoning to the model.

The cost of that inversion is operational. A vector-only memory service is one database. AutoMem is two, and only one of them is optional at runtime. If your recall needs are simple lookups of previously stated preferences, the graph traversal is work you pay for and never use.

A second alternative is doing nothing and relying on long context windows. The README's own BEAM result is the counterargument: at 10 million source tokens, retrieving a few thousand tokens beats feeding the source. Below a few hundred thousand tokens of history, the calculation flips and a long context window is simpler than running two databases.

Licence, maintenance and the upgrade path

AutoMem is MIT licensed. That permits commercial use, modification and redistribution provided the copyright notice and permission notice are included. It does not grant trademark rights, and it carries no warranty. This is not legal advice; if you are embedding the service in a product you ship, have your own counsel read the LICENSE file rather than the summary here.

The repository is not archived, and the last push was on 2026-09-10, which is recent enough that the project is being worked on. Recent releases are v0.16.0 on 2026-06-26, v0.16.1 on 2026-07-07 and v0.16.2 on 2026-08-28. The cadence is roughly monthly, and the patch releases between minors suggest small fixes rather than structural change.

The default branch is develop, not main or master. Anyone pinning a deployment should pin a release tag rather than tracking develop, because the default branch is where work in progress lands.

Upgrade cost is dominated by the two databases, not by the Python code. The Dockerfile builds from python:3.11-slim and installs from requirements.txt, which pins flask 3.0.3, falkordb 1.0.9, qdrant-client 1.11.3 and fastembed 0.4.2, with onnxruntime pinned below 1.20 specifically to avoid issues with that fastembed version. The compose file pins qdrant/qdrant to v1.11.3 but uses falkordb/falkordb:latest, which means FalkorDB can move under you on a rebuild. If you care about reproducibility, tag that image yourself. The compose file also mounts ./backups/falkordb and ./backups/qdrant, so a backup path exists in the layout; the README does not document restore.

Editorial conclusion

Adopt AutoMem if you want memory you host yourself and you are willing to run FalkorDB and Qdrant alongside the Flask API; the graph is the canonical record, so a FalkorDB outage returns 503 and there is no fallback for it. Do not adopt it if a single container is a hard requirement or if you want recall to work with no embedding provider configured. Before committing, verify which embedding provider you will point it at, confirm that the ports in docker-compose.yml (8001 for the API, 6379 for FalkorDB, 6333 for Qdrant) are free on your host, and check INSTALLATION.md for the current setup path rather than relying on the README.

Frequently asked questions

What is AutoMem and what does it do for an AI assistant?

AutoMem is a long-term memory service for AI assistants. It stores decisions, preferences, notes and context in a graph plus a vector store, and returns them to a connected assistant in a later session instead of making you repeat yourself.

How do I install AutoMem?

The repository ships a docker-compose.yml with falkordb, qdrant and flask-api services, so the stack starts with docker compose up -d from the repository root. INSTALLATION.md is the file to check for the current setup path.

Does AutoMem work with Claude and Cursor?

The README says the local MCP bridge supports Claude Desktop, Cursor, Claude Code, Codex and Copilot, and that Remote MCP connects the same service to ChatGPT Developer Mode, Claude.ai and ElevenLabs over HTTPS. The npm package @verygoodplugins/mcp-automem is named as the local bridge.

What happens to AutoMem recall if Qdrant or FalkorDB goes down?

If Qdrant is unavailable the graph still serves recall in a degraded mode. If FalkorDB is down the API returns 503, because the README states the graph is the source of truth.

Official sources

  1. License: MIT
  2. Project website
  3. README
  4. Releases
  5. verygoodplugins/automem on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/verygoodplugins-automem.svg)](https://hysenlabs.com/projects/verygoodplugins-automem)