Model or dataset
MemoriLabs/Memori avatar
MemoriLabs/Memori

Memori: persistent memory for LLM agents, without replacing your stack

Memori is agent-native memory infrastructure. A LLM-agnostic layer that turns agent execution and conversation into structured, persistent state for production systems. Built for enterprise, Memori works with the data infrastructure you already run, no rip-and-replace, and deploys across managed cloud, single-tenant cloud, VPC, and on-premises.

16,779 stars3,483 forksPythonNOASSERTION

At a glance

What is it?
MemoriLabs/Memori is an LLM-agnostic memory layer that turns agent execution and conversation into structured state. Here is how the SDK registers with your client, what the LoCoMo numbers actually say, and where the project's own documentation stops.
Who is it for?
Adopt Memori if you run Python or TypeScript agents against OpenAI-compatible clients and want conversation state persisted without rewriting your prompts or swapping your database, and if you are willing to accept that the SDK's own package metadata still says Alpha.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 26 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Memori actually persists, and who that is for

Most agent frameworks treat memory as a prompt-construction problem. You retrieve some past turns, splice them into the context window, and send the whole thing again. Memori takes a different position, stated in its own tagline: memory from what agents do, not just what they say. The README describes the layer as turning agent execution and conversation into structured, persistent state, and the repository topics list agent-memory, long-short-term-memory, state-management and stateful alongside the usual llm and rag labels.

The audience is narrower than the tagline suggests. This is for teams already running agents in production who have hit the point where conversation history no longer fits, or where the cost of resending it every turn has become the dominant line item. The project targets enterprise deployment explicitly, listing managed cloud, single-tenant cloud, VPC and on-premises as supported shapes, and the README's framing is about fitting the data infrastructure you already run rather than replacing it. If you are prototyping a single-user chatbot with a short session, none of this applies to you yet.

The second audience is teams running specific agent runtimes. The repository ships two named integrations beyond the raw SDK: an OpenClaw plugin and a Hermes Agent memory provider. Both are described as capturing memory from completed turns in the background, with Hermes additionally exposing memori_recall and memori_recall_summary as tools the agent can call itself.

How the SDK hooks into your LLM client

The mechanism is registration, not interception of your prompts. In the Python quickstart, you construct an OpenAI client, then call Memori().llm.register(client) and bind the session to an identity with mem.attribution(entity_id="user_123", process_id="support_agent"). From that point the README states that conversations are persisted and recalled automatically in the background. Your chat.completions.create calls are unchanged.

That design has a consequence worth naming. Memory is scoped by the pair of entity_id and process_id, so the same underlying store can hold separate memory for different end users and different agent roles without you partitioning anything by hand. The TypeScript SDK uses the same two arguments in the same order: .attribution('user_123', 'support_agent').

The retrieval path is not described in the README. It states that Memori recalls the earlier fact, but it does not document how a stored memory is selected for a given turn, what the ranking looks like, or whether you can inspect the candidate set. The benchmark section gives one number that speaks to this indirectly: an average of 721 tokens per query, described as 2.8% of the full-context footprint. That implies retrieval returns a bounded, small payload rather than replaying history, but the selection logic itself is in the docs the README links to, not in the README.

Below the Python surface there is a compiled core. The repository's setup.py declares a RustExtension named memori_python built from core/bindings/python/Cargo.toml with PyO3 bindings, and the pyproject.toml notes that embeddings are provided by the native Rust fastembed core in the default wheel, with the embeddings extra kept as a no-op for compatibility. So pip install memori pulls a binary extension, not pure Python.

Installing Memori and getting a first recall working

Both SDKs are published to their language registries. The README gives the two install commands directly.

bash
pip install memori
bash
npm install @memorilabs/memori

The Python package requires Python 3.10 or newer according to pyproject.toml, and the classifiers list support through 3.14. Note that the build backend requires setuptools-rust, so a source install needs a Rust toolchain; a wheel install from PyPI does not.

The quickstart path needs two credentials. The README instructs you to sign up at app.memorilabs.ai, get a Memori API key, and set MEMORI_API_KEY alongside your LLM key such as OPENAI_API_KEY. Then the Python flow is four lines before your first chat call.

python
from memori import Memori
from openai import OpenAI

client = OpenAI()
mem = Memori().llm.register(client)

mem.attribution(entity_id="user_123", process_id="support_agent")

After that, send a message stating a fact, then send a second message asking about it. The README's example uses "My favorite color is blue." followed by "What's my favorite color?" and states that Memori recalls the answer. What you should see is the second response reflecting the first turn without you having inserted the earlier message into the request yourself.

If you would rather not use the hosted control plane, the README points to a separate BYODB documentation path at memorilabs.ai/docs/memori-byodb/. The repository also carries runnable examples for postgres, mongodb, sqlite, tidb, cockroachdb, neon, oceanbase and others under examples/, which is the most concrete signal of which stores are actually exercised.

The LoCoMo numbers and what they do not cover

The README reports that Memori was evaluated on the LoCoMo benchmark for long-conversation memory and achieved 87% overall accuracy at an average of 721 tokens per query, which it describes as 2.8% of the full-context footprint. It further states that Memori outperformed Zep, LangMem and Mem0 while reducing prompt size by roughly 67% against Zep and lowering context cost by more than 36x versus full-context prompting. A paper is linked at arxiv.org/abs/2603.19935 and results live under docs/memori-cloud/benchmark/.

Treat those as claims from the project's own evaluation rather than independent measurements. The reduction figures are relative to baselines the project chose, and the README does not state the hardware, the model, the number of conversations, or the variance beyond an image caption mentioning standard deviation. If the token reduction is the reason you are evaluating Memori, the benchmark overview and results files are where the methodology would have to be, and they are the first thing to read rather than the headline number.

There is a second gap. LoCoMo measures long-conversation recall. The README's opening claim is broader, covering what agents do, including tool calls, decisions and outcomes. The OpenClaw section mentions capturing tool calls and outcomes explicitly. Nothing in the README reports an evaluation of that execution-level capture, so the benchmark number and the product's central claim are not measuring the same thing.

Where Memori is the wrong choice

The most obvious limitation is the one stated in the package metadata. pyproject.toml carries the classifier Development Status :: 3 - Alpha. That is the project's own label for the Python SDK, and it sits awkwardly next to the enterprise deployment language in the README. If your procurement process requires a stable API guarantee, an Alpha classifier is the fact that matters, not the cloud tier list.

The second limitation is operational. The quickstart depends on a Memori API key issued by app.memorilabs.ai. The README does not document what happens when that control plane is unreachable, whether the SDK buffers locally, or how memory is migrated between the hosted and BYODB paths. It also does not document rollback: there is no described procedure for removing Memori from an application and reconciling the state it wrote. For a layer whose entire job is to hold persistent state, the absence of a documented exit path is a real gap, and it is a gap in the documentation rather than necessarily in the software.

The third is licensing ambiguity. The README badge and the pyproject.toml license field both say Apache-2.0, but the repository's licence identifier is reported as NOASSERTION, meaning GitHub could not classify the LICENSE file. Those two facts are in tension, and the LICENSE file itself is the only thing that resolves it. None of this is legal advice; it is a reason to read the file before you depend on the answer.

Finally, if your memory needs are small and self-contained, a hosted memory layer adds a network dependency and a credential to manage for a problem a local table and a truncation strategy would solve.

Memori against retrieval-based memory systems

The comparison the README draws is against Zep, LangMem and Mem0, all described as retrieval-based memory systems. The stated difference is outcome rather than category: Memori reports lower prompt size at comparable or better accuracy on LoCoMo, roughly 67% less prompt than Zep in the project's measurement.

The architectural distinction that follows from the rest of the README is where the state lives. A retrieval-based system typically embeds and searches a vector store you own, and the memory is whatever the retriever returns. Memori's framing is structured persistent state, with the SDK registering against your LLM client and a compiled core handling embeddings, plus a dashboard at app.memorilabs.ai for browsing memories, analytics, a playground and API keys. The BYODB path exists precisely so that the structured state can sit in a database you already operate rather than one Memori operates for you.

That is a meaningful trade. You get a managed layer with a UI and a bounded token footprint, and in exchange you take on a dependency on a control plane and a schema you did not design. A team that has already built retrieval over its own Postgres with pgvector is not obviously better off here, and the README does not argue that it is. The argument it makes is about token cost and accuracy on long conversations, and that argument only lands if the benchmark methodology holds up for your workload.

Maintenance cadence, upgrade cost and the licence question

The repository is not archived, and the last push was on 2026-09-03, which is recent. That is the whole of what can be said about activity from the available facts. The release history shows a tight run in May 2026: v3.3.4 on 2026-05-20, v3.3.5 on 2026-05-27, and v3.3.6 on 2026-05-28. The version in pyproject.toml is 3.3.7, which is ahead of the newest published release listed, so the main branch carries unreleased work.

Upgrade cost is not documented in the README. There is no migration guide, no statement about whether stored memory survives a version bump, and no compatibility matrix between SDK versions and the cloud API. The CHANGELOG.md file exists at the repository root, and that is where upgrade risk would have to be assessed. For a package that is pre-1.0 in spirit despite a 3.x version number, and labelled Alpha, pinning the version and reading the changelog before each bump is the practical posture.

The dependency footprint is worth noting for upgrade planning. The Python package pulls aiohttp, botocore, faiss-cpu, grpcio, numpy, protobuf constrained below 6.0.0, pyfiglet and requests, plus a Rust extension. The protobuf upper bound in particular is the kind of constraint that collides with other services in a shared environment.

On licence, the README badge and pyproject.toml both declare Apache-2.0, while the repository's licence identifier is reported as NOASSERTION. Apache-2.0 is permissive and carries an explicit patent grant, but the discrepancy means you should read the LICENSE file rather than the badge. Nothing here is legal advice.

Editorial conclusion

Adopt Memori if you run Python or TypeScript agents against OpenAI-compatible clients and want conversation state persisted without rewriting your prompts or swapping your database, and if you are willing to accept that the SDK's own package metadata still says Alpha. Do not adopt it if you need a fully self-contained memory store with no hosted component, since the quickstart path depends on a Memori API key from app.memorilabs.ai and the BYODB route is documented separately. Before committing, verify three things yourself: whether Memori Cloud is the only supported control plane for your deployment shape, whether the BYODB documentation covers the database you already run, and what the actual licence file in the repository says, because the package metadata declares Apache-2.0 while the repository's licence field is reported as NOASSERTION.

Frequently asked questions

What is Memori in the context of LLM agents?

Memori is described in its README as agent-native memory infrastructure, an LLM-agnostic layer that turns agent execution and conversation into structured, persistent state. It installs as a Python or TypeScript SDK and registers against your existing LLM client rather than replacing it.

How do I install Memori?

The README gives two commands: pip install memori for the Python SDK and npm install @memorilabs/memori for the TypeScript SDK. The Python package requires Python 3.10 or newer and ships a compiled Rust extension, so a source build needs a Rust toolchain.

Is Memori the same as the word memori?

No. This Memori is the MemoriLabs/Memori repository, a memory layer for LLM agents written primarily in Python. Searches about the meaning of the word refer to unrelated subjects and have nothing to do with this project.

Official sources

  1. Issues
  2. MemoriLabs/Memori on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/memorilabs-memori.svg)](https://hysenlabs.com/projects/memorilabs-memori)