LongMemory: a local-first memory engine for LLM agents
Local persistent memory store for LLM applications including claude desktop, github copilot, codex, antigravity, etc.
At a glance
- What is it?
- CaviraOSS/LongMemory stores durable, temporal memory in SQLite and exposes it through a TypeScript library, CLI, HTTP server and MCP. The design is opinionated about provenance and time; the documentation is thinner on operational details.
- Who is it for?
- Adopt LongMemory if you need point-in-time truth, provenance and scope enforcement inside a self-hosted process, and you are willing to read Why.md and ARCHITECTURE.md before wiring it into an agent. Do not adopt it if you only need nearest-neighbour chunk retrieval, since the temporal and governance layers are the whole point and they add configuration surface.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 10 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap LongMemory targets: RAG without time or provenance
The README makes its argument explicitly. Most systems called memory, it says, are retrieval pipelines: split text into chunks, embed the chunks, return the nearest vectors. That pipeline answers similarity questions well and answers questions about state badly. It does not tell you what was true in January, whether a newer fact replaced an older one, which source is authoritative, who is allowed to see a record, or why a result deserved a place in the context window.
LongMemory is aimed at teams building agents that accumulate facts over weeks and then have to reason about them. The README frames the target as LLM applications and autonomous agents, with named integrations for Claude Code, Codex, OpenCode, Gemini CLI, Copilot Chat and Cline through a component it calls the session porter. The pitch sentence is that your model stays stateless while your application stops being amnesiac. That is a positioning statement, not a benchmark, and the README does not quantify the difference.
What is concrete is the list of concerns the project models directly: recorded time versus valid time, immutable content and vectors, typed relationships that participate in recall, scope enforcement across project, tenant, user, team, role, agent and task, and a lifecycle with decay, reinforcement, consolidation and reconsolidation. Whether each of those is fully implemented is something the README asserts rather than demonstrates. The design rationale is pushed to a separate Why.md file, which is the document to read before trusting the feature list.
How the hydrograph substrate actually works
The README calls the storage model a hydrograph substrate: immutable nodes, executable edges, worlds, entities, facets and traces. Nodes hold content, vectors, hashes and provenance, and the README states these are not rewritten by recall or by decay. Edges are typed and, in the README's phrasing, executable, meaning they participate in recall and in the explanation of a result rather than sitting as inert metadata.
On top of that graph sit recall modes, which are the part of the design most worth understanding. Strict recall applies temporal, contradiction, contract, confidence and grounding gates. Historical recall preserves superseded truth, so a fact that was later replaced can still be retrieved for the period when it held. Associative recall follows semantic, lexical, entity, activation and graph signals together. World-grounded recall requires current external evidence, which means it can fail when nothing external corroborates the claim.
The gates are configurable through environment variables. The example env file sets LONGMEMORY_STRICT_CONFIDENCE_THRESHOLD to 0.5 and LONGMEMORY_GROUNDING_THRESHOLD to 0.6, and caps assembled context at LONGMEMORY_MAX_CONTEXT_TOKENS of 2048. Those three numbers decide most of the behaviour you will observe in practice. Raising the confidence threshold makes strict recall quieter; lowering the grounding threshold makes world-grounded recall more willing to answer. Nothing in the README explains how the thresholds were chosen, so treat them as defaults to tune rather than as validated values.
Embeddings come from a provider you select. The env file lists openai, gemini, nvidia, aws, ollama, local, siray and synthetic, with three tiers: hybrid and fast use deterministic local embeddings plus lexical ranking, smart mixes 35 percent deterministic with 65 percent from the selected semantic provider, and deep uses the semantic provider directly. The default tier in docker-compose.yml is deep with LONGMEMORY_EMBEDDING_DIMENSION set to 1536, which implies you need a working provider before the container is useful.
Install and first ingest: npm, CLI, then SQLite
The fastest path is the library, which the README says needs no service and no external database for in-memory use. Install it and create a memory instance, ingest a fact, then recall it in strict mode.
npm install longmemoryimport { createMemory } from 'longmemory';
const memory = await createMemory();
await memory.ingest({
user_id: 'alice',
text: 'I prefer TypeScript for backend services',
});
const result = await memory.recall({
text: 'What language does Alice prefer?',
mode: 'strict',
});
console.log(result);
await memory.close();You should see a result object rather than an exception; the README does not print the shape of that object, so log it yourself before parsing fields. When you want the data to survive a restart, switch the store to SQLite and give it a path.
const memory = await createMemory({
store: 'sqlite',
db_path: './longmemory.db',
tenant_id: 'acme',
user_id: 'alice',
});The README states that reopening the same database restores nodes, worlds, entities, edges, temporal history, grounding and lifecycle state. That claim is the one to test first on your own data, because everything else depends on it. There is also a global CLI, which is convenient for checking what the store holds without writing a script.
npm install --global longmemory
longmemory init
longmemory recall "current project priorities" --mode associativeThe package requires Node 20 or newer according to package.json, and the repository pins pnpm 11.5.2 as its package manager. If you build from source instead, the README's sequence is corepack enable, pnpm install --frozen-lockfile, pnpm build, pnpm start, and the API listens on http://127.0.0.1:7331 by default.
Running it as a service, and the MCP surface
For agent hosts, the interesting entry point is MCP rather than the HTTP API. The README gives two forms. A local stdio server bound to a database file and a project name:
longmemory mcp --db .longmemory/project.db --project currentAnd an authenticated Streamable HTTP MCP server, which requires an API key in the environment:
LONGMEMORY_API_KEY=change-me longmemory serve --mcp-httpThe README says the server exposes 13 high-level governed tools plus readable resources and agent workflow prompts, and that tool arguments cannot override the server-bound runtime identity. That last constraint is the security-relevant one: an agent calling a tool cannot widen its own scope by passing a different tenant or user. The README does not enumerate the 13 tools, so you will need to inspect the server or the docs directory to see what is actually callable.
For Docker, the README gives a single-container command and a compose path. The compose file builds from the local Dockerfile, maps port 7331, mounts a named volume at /data, and sets LONGMEMORY_MCP_HTTP to true by default. Note that LONGMEMORY_API_KEY defaults to an empty string in both .env.example and the compose file, and LONGMEMORY_RATE_LIMIT_ENABLED defaults to false. Running the container as-is on a reachable network without setting the key is not something the documentation warns you about.
The Dockerfile runs as a non-root longmemory user, declares /data as a volume, and installs a healthcheck that polls /health every 30 seconds with a 20 second start period. The dashboard is a separate compose profile, started with docker compose --profile ui up --build -d, and serves on http://127.0.0.1:3000. The default allowed origin in .env.example is that same dashboard address.
Where LongMemory is the wrong tool
If your requirement is a vector index over a fixed document set, LongMemory is heavier than the job. You would be running a service, choosing an embedding provider, and tuning confidence and grounding thresholds to get behaviour that a plain nearest-neighbour lookup gives you with less configuration. The README's own framing supports this: the temporal, governance and lifecycle layers are the product, and a workload that does not need them pays for them anyway.
The second limitation is that the correctness story rests on gates with default thresholds that the README does not justify. Strict recall with a confidence threshold of 0.5 will suppress results, and the documentation does not describe how the confidence value is computed or what a reasonable range looks like for a given corpus. If your application cannot tolerate a recall that returns nothing, strict mode is the wrong mode and you should look at associative recall instead.
Third, the lifecycle features are off by default. LONGMEMORY_ENABLE_CONSOLIDATION and LONGMEMORY_ENABLE_COLD_LOG are both set to false in .env.example. Consolidation is one of the capabilities the README lists under lifecycle, so a default deployment does not exercise it. Turning it on is a decision you make without documented guidance on what it does to existing nodes.
Finally, the README describes benchmarks for LongMemEval, LoCoMo and BEAM and the package.json exposes bench, bench:full and bench:ci scripts, but the README does not report results. The harness exists; the numbers are not published there.
LongMemory compared with Mem0 and plain vector stores
The closest comparison in the same space is Mem0, which also targets persistent memory for LLM applications. The difference in approach is where state lives and what the system refuses to do. Mem0-style memory layers typically extract facts from conversation, store them, and return relevant ones, with the extraction and update logic in the pipeline. LongMemory instead keeps content immutable and models supersession as data: the old node is not deleted, it stops being the current truth, and historical recall can still return it. That is a genuine architectural difference, not a labelling one, and it is why valid_time appears as a parameter in the historical recall example.
Against a plain vector database such as a standalone Chroma or Qdrant deployment, the difference is that LongMemory owns the whole path: storage, embedding provider selection, scope enforcement and context assembly under a token cap. You trade the freedom to swap the retrieval layer for a system where permissions and provenance are enforced in one place. Whether that trade is good depends on whether you have multi-tenant or multi-agent scope requirements. If you do not, the governance layer is dead weight.
The honest caveat is maturity. The most recent release listed is v1.3.0, tagged as a beta, from 2025-12-20. The last push to the repository was on 2026-08-31, so the codebase is recent, but a beta tag on the latest release means the API surface may still move between versions. MIGRATION.md exists at the repository root, which suggests the maintainers expect breaking changes to need documentation.
Licence, upgrade cost and what to check before adopting
LongMemory is Apache-2.0, both in the repository metadata and in the package.json license field. That is a permissive licence with an explicit patent grant and a requirement to preserve notices; it does not impose copyleft obligations on your application. This is a description of the licence identifier, not legal advice, and if you are embedding the engine in a distributed product you should read the LICENSE file and the GOVERNANCE.md document at the repository root rather than relying on the identifier alone.
The upgrade cost is the part the documentation leaves open. There is a MIGRATION.md file, and the release history shows v1.2.2, v1.2.3 and v1.3.0 within about two weeks of each other in December 2025, with v1.3.0 marked beta. Rapid patch releases in a short window are normal for a project at this stage, but they also mean you should pin a version rather than track latest, particularly for the Docker image, where the compose file defaults LONGMEMORY_IMAGE_TAG to latest.
The persistence format is the real upgrade risk. The build script copies src/stores/sqlite/schema.sql into the dist output, which means the SQLite schema ships with the package and can change between versions. Before upgrading, back up the database file at LONGMEMORY_DB_PATH, since the README does not document a rollback procedure for a schema change. The README also does not state whether a newer version will open a database written by an older one.
One operational detail worth knowing: .env.example sets LONGMEMORY_TELEMETRY to true and notes that this is local runtime telemetry only, with no outbound telemetry sent. If your environment requires that flag to be off, it is a one-line change, but the default is on.
Editorial conclusion
Adopt LongMemory if you need point-in-time truth, provenance and scope enforcement inside a self-hosted process, and you are willing to read Why.md and ARCHITECTURE.md before wiring it into an agent. Do not adopt it if you only need nearest-neighbour chunk retrieval, since the temporal and governance layers are the whole point and they add configuration surface. Before committing, verify three things yourself: that your embedding provider is reachable from the container (the compose file points Ollama at host.docker.internal:11434), that LONGMEMORY_API_KEY is set to something other than an empty string, and that the recall mode you plan to depend on behaves as documented on your own data.
Frequently asked questions
What is LongMemory and who is it for?
LongMemory is a local-first memory engine for LLM applications and autonomous agents, distributed as a TypeScript package with a CLI, HTTP server, MCP surface and dashboard. It is aimed at teams whose agents accumulate facts over time and need to reason about when those facts were true.
How does LongMemory differ from short-term memory in an agent?
The README draws the line at persistence and time: LongMemory stores content durably in SQLite with separate recorded and valid time, so a fact that was superseded can still be retrieved for the period when it held. A context window holds only what fits in the current prompt.
How do I install and run LongMemory?
The README gives three paths: npm install longmemory for library use, npm install --global longmemory followed by longmemory init for the CLI, and a Docker image at ghcr.io/caviraoss/longmemory that serves the API on port 7331. Building from source uses corepack enable, pnpm install --frozen-lockfile, pnpm build and pnpm start.
What is the difference between strict, historical and associative recall in LongMemory?
Strict recall applies temporal, contradiction, contract, confidence and grounding gates. Historical recall preserves superseded truth and accepts a valid_time parameter. Associative recall follows semantic, lexical, entity, activation and graph signals together.
Does LongMemory require an external database or a cloud service?
The README states that no service or external database is required for in-memory use, and that SQLite persistence is available by setting store to sqlite with a db_path. Embeddings can come from a local provider such as Ollama, which the compose file points at host.docker.internal:11434.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/caviraoss-longmemory)