# LycheeMem: five databases and two disagreeing defaults

> A Python long-term memory framework for agents, shipped as a PyPI package with native plugins for three agent runtimes and an HTTP MCP server for anything else. The page and the environment template disagree on two defaults that matter, the package name is the project's third name, and the test and lint tools are declared as runtime dependencies.

**LycheeMem/LycheeMem** — Lightweight Long-Term Memory for LLM Agents.

- Repository: https://github.com/LycheeMem/LycheeMem
- Website: https://lycheemem.github.io
- Stars: 1,094 · Forks: 18
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/lycheemem-lycheemem

## The project has three names and a sibling with the fourth

Getting oriented here takes a minute, because the name changes depending on where you look.

The PyPI package is `lycheemem`. The repository is `LycheeMem`. The product name used throughout the page is `LycheeMemory`, and a dated news entry records the rename. The manifest is at version 0.1.4 and its description field contains only the old name, with no actual description, which is the kind of leftover that survives a rename because nothing validates it.

The confusing part is the sibling list. The project describes itself as part of a third-generation model series focused on memory intelligence, continual learning and long-context reasoning, and it lists four entries. One of them is a different project that also calls itself LycheeMemory, described as a unified framework for implicit long-term and explicit working memory collaboration, with a paper, a different repository and a seven-billion-parameter model on a model hub.

Within that same list, this project appears under its older name. So the ecosystem contains a paper called LycheeMemory, an infrastructure project called LycheeMem, and a brand that uses both spellings interchangeably. Anyone searching for one of them will land in the wrong place at least once.

## Five integration surfaces over one memory core

The page opens with a row of integration cards, and the claim above them is that it works across agent runtimes that support plugins, MCP, or Python integration.

The five surfaces are distinct in kind, not in branding. One runtime gets a native plugin. One gets an MCP integration plus hooks. A third gets a runtime plugin. The PyPI package is a Python API for direct use. And any MCP client at all is served by an HTTP MCP server.

A dated news entry says what the plugin integrations actually bring: automatic recall, turn mirroring, and a consolidation workflow. Turn mirroring is the piece worth pausing on, because it means the runtime's conversation is being written into memory as it happens, rather than being summarized at the end. That is what makes retrieval work across sessions, and it is also the reason the framework needs its own session database.

A later entry adds OpenAI-compatible chat completions endpoints, with request-level control over consolidation through a consolidate or store parameter. That is the escape hatch for anyone who does not want to drive the memory system through a plugin at all: a chat-shaped API that also decides when to write.

## Five storage paths and two thresholds on one budget

The environment template is where the real architecture shows, and it names five separate storage locations under a data directory.

A sessions database, a skill store database, a skill vector path, a compact memory database, and a compact vector path. The split is legible: relational state for sessions and skills, vector storage for skill retrieval and for the compact semantic memory, and a separate database for that memory's non-vector content. A dated entry explains the design, saying semantic memory was upgraded to use SQLite plus a vector database with no graph database required, which is why the graph library in the dependency list is not a storage engine.

The working memory budget is a single number with two fractions attached: a maximum of 128,000 tokens, a warning threshold at seventy percent, and a block threshold at ninety percent. So the framework has a defined point where it warns you that a conversation is approaching its ceiling and a defined point where it stops, and those two numbers are the entire policy.

Four smaller knobs sit alongside. A minimum recent-turns setting defaults to four, a skill top-k defaults to three, and two boolean switches default to off: one for a composite filter and one for an adequacy check. Both are opt-in, and neither is described in the visible page.

## The reranker is on by default in one place and off in the other

This is the most consequential inconsistency in the repository, because it changes what runs after you install.

The page says the recommended install is the package with its rerank extra, that the extra adds PyTorch and Transformers runtime dependencies, and that with it installed the framework enables a hosted reranker checkpoint by default. It also says that without the extra the core memory system still works and reranking falls back safely.

```bash
pip install lycheemem
```

The variant the page recommends is the same command with the rerank extra appended in quotes, and the backend is then started with the command-line entry point.

The environment template says something different. The reranker is disabled, its backend is set to local rather than hosted, and the model is a specific cross-encoder. So a user who followed the recommended install and then created their configuration file from the template would have a hosted checkpoint enabled by the installer and a local cross-encoder disabled by the file.

The template also shows the shape of a local setup in detail: a device setting that auto-detects, a batch size of sixteen, a maximum sequence length of 512, a composite candidate limit of 80, and a fallback limit of 100. Two commented examples point at running the embedder and the reranker as shared HTTP services on fixed ports, which is how you would put one GPU behind several agent processes.

## Reranking only helps when the memory is already in the pool

The reranker's own description sets a narrow expectation, and it is worth reading before assuming it improves recall.

It improves evidence selection when the correct memory is already in the wider candidate pool. That is reranking, not retrieval: it reorders what search already found and cannot find anything search missed. The dependency on a fixed candidate ceiling in the configuration, a composite limit and a fallback limit, is the same idea expressed as a number.

The evaluation claim is a dated news entry pointing at a separate document. It reports positive hit-at-ten gains on one conversational benchmark and on zero-shot fixtures from three others. The phrasing is careful: positive gains, on fixtures, in a document you have to open to see the numbers. No baseline method and no absolute scores are given on the page.

The other evaluation is a single benchmark comparison against a different runtime's built-in memory, using the plugin integration, reporting a six percent score improvement alongside a seventy-one percent token reduction and a fifty-five percent cost reduction. The token and cost figures are the more useful half of that claim, and they are the kind of number that is a property of the comparison setup rather than of the framework alone.

## The embedding dimension is 1536 in the page and 1024 in the template

The two configuration snippets in the same repository disagree, and the disagreement is the kind that produces a confusing error much later.

The page's snippet sets the embedding model to a hosted small model and its dimension to 1536. The template sets the dimension to 1024 and leaves the model blank. Everything else in the two snippets lines up: a model, a key, an optional base URL.

The provider configuration is otherwise well thought out. Both the language model and the embedder take a provider-qualified model string, and the embedder's key and base URL are separate from the language model's, so you can chat with one vendor and embed with another without a second configuration block. The listed examples include a hosted chat model, a hosted gemini model, and a locally served chat model running under a local runtime, which is the same idea applied to weights on your own machine.

There is also a local embedding path, as an optional extra that adds a sentence-transformers library and PyTorch, and a shared-server path where an embedding server on a local port is addressed as a base URL. So there are four ways to get embeddings, and one number that is wrong in one of the five places you would read it.

## Test and lint tools are declared as runtime dependencies

The manifest closes with a dependency list that should make you look twice.

The test framework, its asyncio plugin, and the linter are all in the runtime dependency list rather than in an optional group. Every install of the package pulls them, and anyone with a locked environment inherits them. Two of the three are also dev-time concerns that a library should not force on its consumers, so this looks like a development dependency list that grew into the production one.

The interpreter range and the tooling target also disagree. The package supports Python 3.9 and newer, while the linter is configured to target 3.11, so code passing the linter can still be using syntax the package claims to support. The build backend is a single modern one, the wheel packages only the source directory, and there is a top-level entry script sitting outside it, which is a layout that works but leaves the question of what that script is for unanswered by the manifest.

The rest of the dependency list is coherent and explains the architecture: a graph library, a tokenizer library for the working memory budget, password hashing and a token library for authentication, a vector database, an image library for the visual memory module, and two model provider clients alongside the provider-agnostic layer that calls them. A dedicated vision directory and a web demo directory sit at the root next to the single example script.

## Conclusion

LycheeMem is a memory layer rather than a model, and its shape suits a team that already has an agent runtime and wants recall, consolidation and a queryable store without building them. Three things to check before wiring it in. The reranker behaves differently depending on how you installed, since the page says the extra turns on a hosted checkpoint by default while the environment template ships it disabled with a local backend selected. The embedder dimension differs between the two configuration snippets in the same repository, so a mismatch is a real possibility rather than a hypothetical. And the reranker only helps when the right memory is already in the candidate pool, which means the retrieval half of the quality story is still whatever your own search does.

## FAQ

### How do I install LycheeMem?

You need Python 3.9 or newer and an API key for an LLM provider. Install the core package with pip install lycheemem, or the recommended variant with pip install "lycheemem[rerank]" to get the transformer reranker and its PyTorch and Transformers dependencies, then start the backend with lycheemem-cli. You can also clone the repository and run pip install -e . to work from source.

### Which LLM providers does LycheeMemory support?

Any provider reachable through litellm, configured as a provider-qualified model string. The page lists an OpenAI model, a gemini model, a locally served Ollama chat model, and any OpenAI-compatible endpoint. The language model and the embedder have separate model, key and base URL settings, so they can point at different providers.

### What does the transformer reranker do in LycheeMemory?

It reranks candidates for semantic memory search and improves evidence selection when the correct memory is already in the wider candidate pool. The page reports positive hit-at-ten gains on LoCoMo and on zero-shot fixtures from LongMemEval-S, MSC-MemFuse and HotpotQA, detailed in a separate document.

### How do I run LycheeMem inside my agent runtime?

There are five documented surfaces: a native OpenClaw plugin, an MCP integration with hooks for Claude Code, a runtime plugin for Hermes, the Python API installed from PyPI, and an HTTP MCP server that any MCP client can connect to. A dated news entry notes the plugin integrations bring automatic recall, turn mirroring and a consolidation workflow.

### What is the OpenAI-compatible endpoint in LycheeMem for?

It is available as of a July 2026 news entry and adds request-level control over consolidation through a consolidate or store parameter, so a client can decide per request whether a turn should be written into memory rather than having the framework decide on its own schedule.

## Sources

- [Issues](https://github.com/LycheeMem/LycheeMem/issues)
- [License: Apache-2.0](https://github.com/LycheeMem/LycheeMem/blob/master/LICENSE)
- [LycheeMem/LycheeMem on GitHub](https://github.com/LycheeMem/LycheeMem)
- [Project website](https://lycheemem.github.io)
- [README](https://github.com/LycheeMem/LycheeMem/blob/master/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/lycheemem-lycheemem
