Model or dataset
aiming-lab/SimpleMem avatar
aiming-lab/SimpleMem

SimpleMem: a self-hosted memory layer for LLM agents, with multimodal retrieval and an MCP server

SimpleMem: Efficient Lifelong Memory for LLM Agents — Text & Multimodal

3,822 stars402 forksPythonMIT

At a glance

What is it?
SimpleMem is an MIT-licensed Python package that stores, compresses and retrieves long-term memory for LLM agents, and exposes text memory to any MCP client. The useful part is the unified package; the part to verify before adopting is whether the paper numbers reproduce in your own data.
Who is it for?
Adopt SimpleMem if you are building an agent that needs cross-session recall and you are willing to run a Python service and point it at an OpenAI-compatible endpoint. Do not adopt it if you need a managed, zero-ops memory API, or if you cannot accept that the retrieval stack pulls in LanceDB, sentence-transformers and tantivy.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 68 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem SimpleMem targets: memory that outlives a single chat session

Most agent frameworks treat context as something you assemble per request. You retrieve a few chunks from a vector store, paste them into the prompt, and throw the whole thing away when the request ends. That works for a single task. It falls apart when the agent is supposed to remember what a user said last week, or what it already tried and failed at.

SimpleMem is aimed at that second case. The README describes it as "Efficient Lifelong Memory for LLM Agents", and the package name on PyPI is `simplemem`. The intended user is an engineer building an agent that accumulates state over many sessions: a coding assistant that remembers project conventions, a support agent that remembers a customer's history, or a research agent that keeps notes across runs.

The project also ships an MCP server, which is the lower-effort integration path. Any MCP client (the README lists Claude Desktop, Cursor, LM Studio and Cherry Studio) can talk to it, so you get persistent text memory without writing retrieval code yourself. Multimodal memory, covering image, audio and video, is only available through the Python integration, not through MCP.

How the memory pipeline is put together

The repository layout tells you most of the architecture. There is a `simplemem/` package, a `simplemem_router.py` at the top level, and separate directories for `OmniSimpleMem/` (multimodal) and `EvolveMem/` (the self-evolution loop). The v0.3.0 release note says these were merged into one package: `from simplemem import SimpleMem` auto-selects the text or multimodal backend from the first method you call.

Storage is LanceDB. The `setup.py` install list pins `lancedb>=0.4.0` and notes that `tantivy>=0.20.0` is "required at runtime by lancedb FTS (use_tantivy=True)", with `pyarrow>=12.0.0` imported at the top of `simplemem/core/database/vector_store.py`. That combination means the store is doing both vector search and full-text search, which is what a hybrid retriever needs. `simplemem/core/hybrid_retriever.py` imports `dateparser`, so time expressions in queries are parsed rather than treated as opaque strings.

Embeddings come from `sentence-transformers`, and generation comes through the `openai` client, which is why the README can point `OPENAI_BASE_URL` at any OpenAI-compatible endpoint. The Docker compose file defaults `LLM_PROVIDER` to `openrouter` and `EMBEDDING_MODEL` to `qwen3-embedding:4b` at `EMBEDDING_DIMENSION=2560`, and it ships an `OLLAMA_BASE_URL` pointing at `host.docker.internal:11434/v1` for a local model on the host. Those defaults are a reasonable hint about the intended deployment shape: a container, a host-side inference endpoint, and a persistent volume.

The compression claim in the README is "semantic lossless compression". The repository does not document the compression ratio or the algorithm, and neither the release notes nor the setup script state what is discarded. Treat that phrase as a design goal rather than a measured property until you can check it against your own corpus.

Installing SimpleMem and running a first retrieval

The README's install path is a local editable install. The v0.3.0 note gives `pip install -e .` as the one-step command for the unified package, and the repository also carries a `requirements.txt` with pinned versions and a `requirements-gpu.txt` for a GPU environment.

bash
pip install -e .

After that, the import is a single line. According to the release note, the backend is chosen from the first method you call, so the same import serves both the text and the multimodal path.

python
from simplemem import SimpleMem

The repository includes `examples/quickstart.py` as the worked example. The README does not reproduce its contents, so read that file in the checkout rather than guessing at constructor arguments.

If you want the MCP server instead of the library, the Docker path is documented through compose. The service maps port 8000 and keeps its data in a named volume, with the compose file explaining that a bind mount to `./data` often has the wrong host ownership for the container user.

bash
docker compose up -d

The compose file also shows the environment variables the server reads: `DATA_DIR`, `LANCEDB_PATH`, `USER_DB_PATH`, `LLM_PROVIDER`, `LLM_MODEL`, `EMBEDDING_MODEL` and `EMBEDDING_DIMENSION`. Note the two secrets it defaults: `JWT_SECRET_KEY` and `ENCRYPTION_KEY`. The defaults are literal placeholder strings, so set both before exposing the port.

Where SimpleMem is the wrong choice

The dependency footprint is the first real constraint. A default install pulls LanceDB, sentence-transformers, tantivy, rank_bm25 and the OpenAI client, and the setup script notes that torch and transformers arrive transitively through sentence-transformers. That is a heavy install for what is, at the API surface, a memory store. If your deployment target is a small serverless function, this is not a drop-in.

The second constraint is that the strongest numbers in the README are paper numbers, not deployment numbers. The release notes claim "new SOTA on LoCoMo (F1=0.613, +47%)" and "Mem-Gallery (F1=0.810, +51%)", and EvolveMem is described as outperforming the strongest baseline by "+25.7% relative" on LoCoMo and "+18.9% relative" on MemBench. Those are relative gains against named baselines on specific benchmarks. Nothing in the repository states what happens on short conversations, on non-English text, or on a corpus where the same fact is restated with contradictions. Retrieval quality on your data is the thing you have to measure, and the repository ships `test_locomo10.py` and a `test_ref/` directory, which suggests the benchmark harness is in-tree but says nothing about your workload.

The third is scope. MCP gives you text memory only. If your agent needs to recall a screenshot or a call recording, you are writing Python against the library, not configuring a client.

Finally, maintenance: the last push to the default branch was on 2026-07-24, and the most recent tagged release is v0.3.0 from 2026-05-21. The repository is not archived, but the gap between the last push and the release tag is worth noting if you depend on fixes landing quickly.

How SimpleMem differs from LangMem and a plain vector store

The closest comparison in the repository is LangMem, which appears in the pinned `requirements.txt` at version 0.0.30 alongside `langgraph` and `langchain`. LangMem is a memory library for the LangChain and LangGraph stack. If your agent already runs on LangGraph, LangMem is the path of least resistance: it lives inside the graph you already have, and you do not stand up a separate service.

SimpleMem's difference is that it is not coupled to an agent framework. It is a package plus an HTTP server, and the MCP endpoint means a client that knows nothing about Python can use it. The trade is that you own the process, the volume and the embedding model. There is no graph to hang it from, and no framework to inherit retries or checkpointing from.

Against a plain vector store, the difference is the retriever. A bare LanceDB or Qdrant collection gives you nearest-neighbour search. SimpleMem's `hybrid_retriever.py` combines vector search with BM25 keyword retrieval (`rank_bm25` is in the install list) and parses date expressions with `dateparser`, so a query like "what did we decide before the migration" can be filtered on time rather than matched purely by embedding similarity. If your queries are all short and topical, a plain vector store is simpler and cheaper. The hybrid path earns its dependencies when your queries carry time references or exact identifiers that embeddings blur.

Licence, upgrade cost and what v0.3.0 changed

SimpleMem is MIT-licensed, and the LICENSE file is at the repository root. MIT is permissive: you can use it commercially, modify it and redistribute it, provided the copyright notice and permission notice travel with it. That is a statement about the licence text, not legal advice, and it does not cover the models you point it at. If you route through a hosted OpenAI-compatible provider, or download `qwen3-embedding:4b` for a local Ollama instance, those terms are separate and are yours to check.

The upgrade story is the part to plan for. v0.3.0 merged three previously separate things (SimpleMem, Omni-SimpleMem and EvolveMem) into one package, and the release note describes auto-routing based on the first method called. That is a structural change, not a patch, and code written against v0.1.0 or v0.2.0 import paths may not survive it. The `simplemem_router.py` file at the top level is a visible artifact of that merge.

On the storage side, the compose file keeps LanceDB data in a named volume at `/app/MCP/data`. Upgrading the image means that volume persists, but the repository does not document a schema migration step. Back up the volume before a major version bump, and check the release notes for that version rather than assuming the store is forward-compatible.

Editorial conclusion

Adopt SimpleMem if you are building an agent that needs cross-session recall and you are willing to run a Python service and point it at an OpenAI-compatible endpoint. Do not adopt it if you need a managed, zero-ops memory API, or if you cannot accept that the retrieval stack pulls in LanceDB, sentence-transformers and tantivy. Before committing, verify three things: that `pip install -e .` resolves cleanly on your Python version, that the MCP server starts on port 8000 with your own `LLM_PROVIDER` and `EMBEDDING_MODEL`, and that the LoCoMo and Mem-Gallery numbers in the release notes hold on a sample of your own conversations.

Frequently asked questions

What is SimpleMem?

SimpleMem is an MIT-licensed Python package from aiming-lab for storing, compressing and retrieving long-term memory for LLM agents. It ships a text-memory MCP server and a Python integration that adds image, audio and video memory.

Is SimpleMem an efficient lifelong memory framework for LLM agents?

That is how the README describes it: "Efficient Lifelong Memory for LLM Agents". The claim rests on retrieval quality on benchmarks such as LoCoMo and Mem-Gallery, so whether it is efficient for your workload depends on your own corpus.

What is short-term memory in AI agents?

The repository does not define short-term memory; it addresses the long-term case, describing memory that persists across sessions through a LanceDB store with hybrid retrieval. For a definition of short-term memory you need a source outside this project.

Official sources

  1. aiming-lab/SimpleMem on GitHub
  2. Issues
  3. License: MIT
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/aiming-lab-simplemem.svg)](https://hysenlabs.com/projects/aiming-lab-simplemem)