doobidoo/mcp-memory-service: a self-hosted memory backend for agent pipelines
Open-source persistent memory for AI agent pipelines (LangGraph, CrewAI, AutoGen) and Claude. REST API + knowledge graph + autonomous consolidation.
At a glance
- What is it?
- One service exposes memory over REST, MCP and a CLI, with a knowledge graph and local ONNX embeddings. The trade-off is that you now run and secure a stateful server.
- Who is it for?
- Adopt it if you already run multi-agent pipelines that need shared, inspectable memory and you are willing to operate a stateful HTTP service, including the anonymous-access flag or OAuth in front of it. Skip it if a single assistant session is your whole use case, since the built-in file memory in most clients covers that with no server to run.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The amnesia problem this service targets
A chat session ends and the reasoning behind it disappears. The README frames the cost bluntly: you re-explain your architecture at the start of the next session. For a single assistant that is annoying. For an agent pipeline it is worse, because the researcher agent, the coder agent and the reviewer agent each start from an empty context and cannot see what the others concluded.
The project's answer is a memory service that sits outside any one agent runtime. Memories are stored once and retrieved by any client that can reach the service, whether that client is Claude Desktop over MCP, a LangGraph node over HTTP, or a shell script over the CLI. The README positions this against the alternative of wiring Redis plus a hosted vector database plus glue code yourself, and against cloud memory APIs that bill per call.
The audience is therefore narrow but real: people building agentic systems, not people using a chat assistant casually. The topics list confirms it, naming LangGraph, CrewAI and AutoGen alongside the MCP angle.
REST API, MCP transport and a knowledge graph in one process
The service is a Python package that can run in two modes. In stdio mode it speaks MCP and is launched by an MCP client such as Claude Desktop. With the --http flag it starts an HTTP server, which the README says listens on http://localhost:8000 and exposes 76 endpoints, so an agent does not need an MCP client library at all.
Storage is local. Embeddings are computed with ONNX on the machine running the service, which means memory content does not leave your infrastructure. The vector storage topic points at sqlite-vec, so the default path appears to be a SQLite file with a vector extension rather than an external database. The repository also carries Cloudflare configuration in .env.example, with CLOUDFLARE_API_TOKEN, CLOUDFLARE_ACCOUNT_ID and a D1 database ID, so a Workers-hosted deployment is supported as a separate path.
Two headers carry the multi-agent semantics. X-Agent-ID auto-tags a stored memory with agent:<id>, and the README's example shows a memory posted by the researcher agent coming back tagged agent:researcher. A conversation_id field bypasses deduplication, which matters when you are appending turn by turn and near-identical sentences would otherwise collapse into one record. The knowledge graph adds typed edges between memories, and the README names causes, fixes and contradicts as edge types, so retrieval can follow a causal chain instead of matching text alone. SSE events notify connected clients when any agent stores or deletes a memory.
Install and store your first shared memory
The package installs from PyPI and requires Python 3.10 or newer according to pyproject.toml. The README's quick start is a single pip command.
pip install mcp-memory-serviceFor Claude Desktop you then add an entry to the client config file. On Linux that file is ~/.config/Claude/claude_desktop_config.json; macOS and Windows paths are listed in the README. The server command is memory with the server argument.
{
"mcpServers": {
"memory": {
"command": "memory",
"args": ["server"]
}
}
}After restarting the client, the memory tools appear. For an agent pipeline you want the HTTP mode instead, started with anonymous access enabled for local experimentation.
MCP_ALLOW_ANONYMOUS_ACCESS=true memory server --httpThe reader should see the REST API at http://localhost:8000. From there a plain httpx client can write and read memory, and the X-Agent-ID header is what scopes a record to one agent.
import asyncio
import httpx
async def main():
async with httpx.AsyncClient() as client:
await client.post("http://localhost:8000/api/memories", json={
"content": "API rate limit is 100 req/min",
"tags": ["api", "limits"],
}, headers={"X-Agent-ID": "researcher"})
asyncio.run(main())The stored record comes back tagged with api, limits and agent:researcher, and a later search filtered on agent:researcher returns it. Note that MCP_ALLOW_ANONYMOUS_ACCESS is exactly the kind of flag you should not leave set on a host reachable from anywhere but your own machine.
Where the design puts work on you
The service is stateful, and that is the main cost. You now own a process with a database behind it, a port to expose or firewall, and a backup question the README does not answer. Every agent that needs memory gains a network dependency on that process; if it is down, retrieval fails and the agent silently loses context unless you handle the error.
Authentication is the sharp edge. The documented local path uses MCP_ALLOW_ANONYMOUS_ACCESS=true, which is fine on localhost and not fine anywhere else. The project also documents OAuth 2.0 with dynamic client registration for remote MCP, and the badge list links an OAuth setup document, but the README does not spell out the threat model for a shared deployment. If several teams point agents at one instance, tag scoping via X-Agent-ID is a naming convention, not an access control boundary. Nothing in the README suggests a memory written by one agent is hidden from another that searches without a tag filter.
The 5ms retrieval figure in the README is a project claim, not an independently verified number, and it will depend on corpus size and hardware. Treat it as a target rather than a specification.
This is also the wrong tool when your problem is a single long conversation. If you only need one assistant to remember a project, the built-in memory of the client is less machinery for the same result.
How it differs from a memory MCP server that is just a file
The obvious alternative is a lightweight memory MCP server that writes Markdown or JSON to a directory and lets the model read it back. Those are trivial to run, easy to inspect in a text editor, and easy to version in git. They also have no embeddings, so retrieval is whatever the model chooses to open, and no shared state, so two agents on two machines do not see the same memory.
mcp-memory-service takes the opposite position on every one of those points. Semantic search over embeddings, one server shared by many clients, a knowledge graph with typed edges, and consolidation that compresses older memories. The repository layout reflects the weight: a src tree, a tests tree, Docker files, a Cloudflare deployment path, OAuth documents and a site directory. You are adopting a service, not a script.
The middle option is a hosted memory API. It removes the operational burden and adds a per-call bill plus the requirement that your agents' context leaves your network. The README's stated pitch is precisely that it avoids both, with embeddings running locally and no cloud cost. That claim holds only if you count your own server time as free, which is the honest way to read it.
Releases, licence and the upgrade question
Versioning is fast. The releases list shows v11.9.0, v11.10.0 and v11.11.0 all dated 2026-09-05, and pyproject.toml in the repository already declares 11.12.0. The last push to the default branch was on 2026-09-10. Three minor releases in one day suggests either a batch of small fixes or a release pipeline that publishes often; either way, pinning a version is the sane default for a service holding your agents' memory.
The README does not document a rollback procedure, and the CHANGELOG is the only release history the repository offers. Since the storage layer is a local database, an upgrade that changes the schema is the scenario worth testing on a copy before you touch the instance your agents depend on.
Licensing is Apache-2.0, confirmed by both the LICENSE file and the classifier in pyproject.toml. That permits commercial use and modification, and it includes an explicit patent grant, which matters if you embed the service in a product. The repository also contains a NOTICE file, which Apache-2.0 expects you to carry forward in redistributions. This is a description of the licence text, not legal advice; your counsel decides what your distribution requires.
Editorial conclusion
Adopt it if you already run multi-agent pipelines that need shared, inspectable memory and you are willing to operate a stateful HTTP service, including the anonymous-access flag or OAuth in front of it. Skip it if a single assistant session is your whole use case, since the built-in file memory in most clients covers that with no server to run. Before committing, verify on your own hardware that the ONNX embedding path loads on your CPU, that sqlite-vec is available for your Python build, and that your chosen auth mode actually rejects unauthenticated writes. Also check the CHANGELOG against your pinned version, because the project ships releases frequently and there is no documented rollback path in the README.
Frequently asked questions
What is an MCP service?
MCP is the Model Context Protocol, and in this project it is one of several transports the memory service speaks. The package can run in stdio mode so an MCP client such as Claude Desktop launches it as a server, or in HTTP mode so any client can reach it over the REST API.
Why would I use an MCP server like mcp-memory-service?
The README's case is that each agent run otherwise starts from zero and you re-explain your architecture every session. This service stores decisions and context outside any one run so multiple agents and sessions retrieve the same memory, with embeddings computed locally rather than through a cloud API.
How do I disable an MCP server such as mcp-memory-service?
The README does not document a disable procedure. The service is registered in the client's MCP configuration, so the practical step is removing its entry from that config file, for example the memory block in claude_desktop_config.json, and restarting the client.
What is open memory MCP in relation to mcp-memory-service?
The project describes itself as open source under Apache-2.0, with the service self-hosted and embeddings run locally via ONNX so memory does not leave your infrastructure. The repository also supports a Cloudflare Workers deployment path configured through .env.example.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/doobidoo-mcp-memory-service)