Open-source project
supermemoryai/supermemory avatar
supermemoryai/supermemory

Supermemory: a self-hostable memory layer for AI assistants and agents

Memory and context engine + app that is extremely fast, scalable, and can be run fully locally. The Memory API for the AI era.

30,945 stars2,712 forksTypeScriptMIT

At a glance

What is it?
Supermemory is an MIT-licensed memory and context engine from supermemoryai, shipped as an API, MCP server and plugins. Here is what the repository actually documents, how to run it locally, and where it stops being the right tool.
Who is it for?
Adopt Supermemory if you want conversation memory and user profiles behind one API and are willing to run it yourself or depend on the hosted service. Do not adopt it if you need a documented rollback path, a schema migration guide or a stable local server; the README covers none of these and the server package is still at 0.0.8.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 4 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

The forgetting problem Supermemory is built around

Every conversation with an AI assistant starts from nothing. You explain your stack, your naming conventions, the client you are mid-project with, and the next session you explain it again. Supermemory targets that gap. The README frames it plainly: "Your AI forgets everything between conversations. Supermemory fixes that."

The project has two audiences and the README splits them explicitly. The first is people who use AI tools: Claude Code, Cursor, Codex, OpenCode, OpenClaw, Hermes and Claude Desktop, reached through plugins or an MCP server. The second is developers building agents and apps who want memory, retrieval and user profiles behind one API instead of wiring a vector database, an embedding pipeline and a chunking strategy themselves. The README lists all three of those as things you do not have to configure.

That second audience is where the scope gets interesting. Supermemory is not just a vector store with a search endpoint. It extracts facts from conversations, maintains user profiles, and according to the README handles "temporal changes, contradictions, and automatic forgetting." Contradiction handling is the part most retrieval stacks leave to the caller: if a user says they moved from Postgres to SQLite, a plain embedding index returns both statements with similar scores. Whether Supermemory resolves that well is something the README asserts rather than demonstrates.

How memory, profiles and hybrid search fit together

The README describes one memory structure and ontology holding everything: extracted facts, user profiles, knowledge base documents and connector content. Three surface areas sit on top of it.

The memory layer pulls facts out of conversations and manages their lifecycle. The profile layer is described as auto-maintained context combining stable facts with recent activity, returned in one call that the README puts at roughly 50ms. Hybrid search is the third piece: RAG over your documents and personalised memory in a single query, so a knowledge base answer and a user-specific answer come back together rather than from two systems you reconcile yourself.

On the client side, the mechanism is three MCP tools. memory saves or forgets information and the assistant calls it when something looks worth keeping. recall searches memories by query and returns relevant memories plus a profile summary. context injects the full profile at the start of a conversation, and the README notes that in Cursor and Claude Code you can trigger it by typing /context. Scope is handled with projects, which the README calls container tags, so work and personal context, or one client and another, stay separated.

The repository layout backs the description. The topics list includes cloudflare-workers, cloudflare-kv, cloudflare-pages, drizzle-orm, postgres, remix, tailwindcss, vite and typescript, and the top level holds apps/, packages/, skills/ and a turbo.json, so this is a Turborepo monorepo with a web app, server packages and a skills directory. The root package.json pins [email protected] as the package manager and requires node >=20. Dependencies include drizzle-orm with pg and postgres, better-auth, hono-openapi and zod, which tells you the server is a TypeScript API over Postgres with schema defined in Drizzle and validation in Zod. That is a conventional stack, and it means self-hosting is a database plus a Node process, not a bespoke runtime.

Running Supermemory locally with one command

The README offers a local path under the heading "Supermemory local: run it yourself," described as one binary and zero config, with the option to bring any model or run fully offline with Ollama. The install is a single shell command.

bash
curl -fsSL https://supermemory.ai/install | bash

Piping a remote script into bash is a decision worth pausing on. The README does not publish a checksum or a signature for that script, so if you work somewhere that requires artifact verification, download it, read it, then run it rather than piping it blind.

For the MCP route there is nothing to install at all. The README gives the server URL and the client config block, which you paste into your MCP client configuration.

json
{
  "mcpServers": {
    "supermemory": {
      "url": "https://mcp.supermemory.ai/mcp"
    }
  }
}

After restarting the client, the three tools (memory, recall, context) should be available. In Cursor and Claude Code the README says typing /context injects your profile.

If you are working on the repository itself rather than consuming the binary, the root package.json defines the scripts. The workspace uses bun, and the relevant entry points are:

bash
bun install
bun run dev

The README does not document what bun run dev:local does differently from bun run dev, even though both exist in the root scripts. That is a gap you will have to resolve by reading the app packages.

Where Supermemory is the wrong choice

The local server is at server-v0.0.8, released on 2026-08-17, with server-v0.0.7 two days earlier. A version below 0.1 across three releases in a month is a young surface, and the README does not describe a schema migration story, a rollback procedure or a data export format. If your memory store becomes load-bearing, you are betting on a project that has not yet published how it handles an upgrade that changes the on-disk shape. The README is silent on all three points.

The second limitation is model and provider coupling. The root dependencies pull in the Anthropic, OpenAI, Google, Cerebras and Vercel AI SDK packages. The claim that you can "bring any model" is not backed in the README by a list of supported providers or a configuration example showing how a custom endpoint is wired. If you run an internal model behind a gateway, assume you will be reading the server source before you know whether it fits.

The third is the split between hosted and local. The MCP server URL points at mcp.supermemory.ai, a hosted endpoint, while the install script and the self-hosting docs describe a local deployment. The README does not state whether the MCP server can be pointed at your own local instance, or what data leaves your machine when you use the hosted URL. For anything touching client work or personal context, that is the question to answer before the first conversation, not after.

Finally, the benchmark claims. The README states first place on LongMemEval, LoCoMo and ConvoMem, with 95% Recall@15 at 99.4% context reduction. Those are the project's own numbers on benchmarks it links to, and the README does not describe the evaluation harness or the configuration used. Treat them as a pointer to the research page, not as a prediction of your workload.

Supermemory against a hand-rolled retrieval stack

The natural alternative is not another memory product. It is the stack most teams already have: a vector database, an embedding model, a chunking step and a prompt that stuffs retrieved text into context. The difference in approach is where the intelligence sits.

In a hand-rolled stack, storage is dumb and the application decides what matters. You chunk documents, embed them, retrieve top-k by similarity, and the model sees whatever came back. Contradictions, recency and per-user personalisation are your problem. You write the logic that notices a user changed their mind, and the logic that decides which of two conflicting facts to surface.

Supermemory moves that logic into the engine. Facts are extracted at write time rather than retrieved raw at read time, and the engine owns temporal change, contradiction handling and forgetting. The README's 99.4% context reduction figure is the practical consequence: instead of retrieving passages, you retrieve distilled facts and a profile, which is a much smaller prompt. The cost is that you no longer control what was extracted. If the extraction drops something your application needed, you cannot fix it by changing a retrieval parameter; you have to change how you write to the memory layer.

That trade is reasonable for conversational memory, where the raw transcript is noise and the distilled fact is the signal. It is a worse fit for document question answering over a fixed corpus, where you usually want the original passage verbatim so the model can quote it. The README's hybrid search is aimed at the case where you need both, and that is the configuration worth testing first.

Licence, maintenance and what an upgrade costs you

The repository is MIT licensed, which permits commercial use, modification and redistribution provided the copyright notice and permission notice are retained. The README does not describe a separate licence for the hosted service, the plugins or the MCP server, and the plugins live in their own repositories under the supermemoryai organisation. If you plan to depend on the hosted API rather than self-host, check the terms attached to that service separately; the MIT grant covers the code in this repository, not a hosted endpoint. None of this is legal advice, and a licence file is not a substitute for reading it.

On maintenance, the signals are current. The repository is not archived and the last push was on 2026-09-17, the same day as the most recent activity recorded here. Releases are frequent: server-v0.0.7-rc.2 on 2026-07-22, server-v0.0.7 on 2026-08-15, server-v0.0.8 on 2026-08-17. Frequent releases at the 0.0.x level usually mean the interface is still moving, which is the upgrade cost you should budget for rather than the licence cost.

The practical upgrade burden sits in the Drizzle schema. Because persistence goes through drizzle-orm against Postgres, a schema change in a new server release means a migration on your database, and the README does not document how those migrations are generated or applied. If you self-host, pin the server version, keep a database snapshot before upgrading, and read the release notes for that tag before you pull. The root package.json also pins [email protected] and requires node >=20, so a Node 18 host will fail before it reaches any memory logic.

Editorial conclusion

Adopt Supermemory if you want conversation memory and user profiles behind one API and are willing to run it yourself or depend on the hosted service. Do not adopt it if you need a documented rollback path, a schema migration guide or a stable local server; the README covers none of these and the server package is still at 0.0.8. Before committing, verify two things: that your target client is on the supported list and that your machine satisfies the package.json engine constraint of Node 20 or newer, plus bun 1.3.6 as the pinned package manager.

Frequently asked questions

What is Supermemory?

It is a memory and context engine for AI, described in the README as the memory and context layer for AI. It extracts facts from conversations, maintains user profiles and serves retrieval through an API, an MCP server and editor plugins.

Can Supermemory be self-hosted?

Yes. The README has a section titled "Supermemory local: run it yourself," described as one binary and zero config, and the install is a single curl command. It also states you can bring any model or run fully offline with Ollama.

Is Supermemory open source?

The repository is MIT licensed and the README says the plugins are open source implementations of the Supermemory API, with separate repositories for Claude Code, Cursor, Codex, OpenClaw and OpenCode. The MCP server is also described as open source, with a link to its source.

How do you use Supermemory with an MCP client?

Add the server URL https://mcp.supermemory.ai/mcp to your MCP client config under the key supermemory. Once connected, the client gets three tools: memory, recall and context, and in Cursor and Claude Code the README says typing /context injects your profile.

Is Supermemory free?

The README does not state pricing. The code in this repository is MIT licensed and the local install path is documented, but the README does not describe the terms of the hosted API or the hosted MCP endpoint.

Official sources

  1. License: MIT
  2. Project website
  3. README
  4. Releases
  5. supermemoryai/supermemory on GitHub
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/supermemoryai-supermemory.svg)](https://hysenlabs.com/projects/supermemoryai-supermemory)
Community notes

Community notes