LightMem: A Lightweight Memory Framework for LLMs and AI Agents
[ICLR 2026] LightMem: Lightweight and Efficient Memory-Augmented Generation
At a glance
- What is it?
- LightMem is a Python framework from the zjunlp group, accepted at ICLR 2026, that adds long-term memory to LLM applications and AI agents. It provides a storage, retrieval, and update layer that works with both cloud APIs and local models, and ships with an MCP server for agent integration.
- Who is it for?
- LightMem is a good choice for researchers evaluating memory augmentation strategies and for developers building conversational agents that need persistent context beyond a single session. It is not suited for production deployments requiring strict version stability: the package is at version 0.1.0, Python is constrained to 3.10 or 3.11, and many dependencies are pinned to specific versions.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 25 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The Long-Context Memory Problem LightMem Addresses
LLM agents lose context between sessions because they have no mechanism to store and retrieve information from past interactions. Within a single session, the context window eventually fills, and relevant earlier content is dropped. LightMem addresses this by providing an explicit memory layer: a modular system for storing conversation history, retrieving relevant memory entries for each new query, and updating or expiring memories over time.
The primary intended use is building AI applications that require long-term memory: conversational assistants that remember user preferences and facts across sessions, research agents that accumulate findings over multiple runs, and multi-agent systems where agents need to share or access a common memory store. The README describes the framework as targeting both LLM developers building intelligent applications and AI safety researchers benchmarking memory strategies on standard datasets.
Memory Storage, Retrieval, and Update Architecture
LightMem's architecture separates memory into three operations: storage, retrieval, and update. Storage handles writing new memories derived from conversation turns. Retrieval selects relevant memories for a given query, using a combination of retrieval strategies that can include BM25 keyword matching, vector similarity (via `qdrant-client`), and LLM-based compression with `llmlingua`.
The framework supports a modular storage backend design: the README describes custom storage engines and retrieval strategies as supported through the modular architecture. The pyproject.toml lists `qdrant-client==1.15.1` as a dependency, indicating vector-based memory retrieval is the primary backend for semantic search.
The LongMemEval evaluation specifically tests offline memory update, meaning the system updates stored memories based on conversation history before answering questions. This mode is documented in the `experiments/longmemeval/readme.md` reproduction script and is distinct from online update, where memories are updated incrementally during a live conversation session.
The repository also includes StructMem, a hierarchical memory variant that preserves event-level memory bindings and cross-event connections (accepted at ACL 2026), FluxMem, a connectivity-evolving memory framework that models memory as a heterogeneous graph (under review at EMNLP 2026), and EM²Mem, an event-centric multimodal memory framework accepted at EMNLP 2026. Each is documented in its own markdown file in the repository root.
Installing LightMem and Running the Baseline Evaluation
LightMem requires Python 3.10 or 3.11 (the pyproject.toml specifies `>=3.10,<3.12`). The package is available on PyPI.
pip install lightmemFor adding DeepSeek model support, the README notes the configuration file at `src/lightmem/configs/memory_manager/base_config.py`, which accepts `deepseek-v4-flash` and `deepseek-v4-pro` with `reasoning_effort` and thinking-mode parameters (added April 2026).
For local model support via Ollama or vLLM, the factory implementations are at `src/lightmem/factory/memory_manager/ollama.py` and `src/lightmem/factory/memory_manager/vllm_offline.py`. A Transformers auto-loading path is at `src/lightmem/factory/memory_manager/transformers.py`.
The baseline evaluation framework at `src/lightmem/memory_toolkits/readme.md` supports running Mem0, A-MEM, and LangMem on LoCoMo and LongMemEval for comparison. Reproduction scripts for LightMem's own results on LoCoMo are at `experiments/locomo/readme.md` and for LongMemEval at `experiments/longmemeval/readme.md`. Tutorial notebooks in `tutorial-notebooks/` cover common usage scenarios.
MCP Server and Agent Integration
LightMem ships an MCP server at `mcp/server.py`, added in November 2025. This makes LightMem's memory tools available to MCP-compatible agents, including coding agents that support the Model Context Protocol.
The MCP server exposes LightMem's operations (read, write, update, retrieve) as tools that an agent can call during its execution. This integration path is relevant for teams building agents with Claude Code, OpenClaw, or other MCP-compatible orchestrators that need persistent memory without managing a separate memory service.
The repository also includes a `web/` directory, indicating a web interface for interacting with the memory system. The tutorial notebooks in `tutorial-notebooks/` cover multiple usage scenarios and were released alongside a demo video in December 2025. The CCF ODTC open source incentive program selected LightMem in July 2026, which indicates the project has received recognition from the Chinese open-source community beyond its ICLR acceptance.
Benchmarks: LoCoMo and LongMemEval
LightMem includes evaluation results on two long-term memory benchmarks. LoCoMo tests memory retrieval over long social media conversation threads. LongMemEval tests both retrieval and offline memory update on long question-answering tasks.
The repository provides scripts for reproducing both sets of results, along with a unified baseline framework that runs Mem0, A-MEM, and LangMem on the same datasets for comparison. The README describes leading results on LoCoMo with strong performance and efficiency, directing readers to the reproduction scripts for the specific numbers.
The zjunlp group maintains a separate companion repository, MemBase, which provides a more comprehensive baseline evaluation framework supporting additional memory systems (including EverMemOS) on both LoCoMo and LongMemEval. This companion repository was announced in March 2026 and extends the comparison surface beyond what is included in the LightMem repository itself.
The dependency pin strategy (torch==2.8.0, transformers==4.57.0, specific library versions throughout) reflects the typical research artifact approach of locking dependencies for reproducibility rather than maintaining forward compatibility.
Scope Boundaries and Alternatives
LightMem is a research framework. Version 0.1.0 and the tight Python version constraint (3.10 or 3.11 only) signal that the project prioritizes reproducibility and research flexibility over production deployment stability. Teams building production agents that need a supported, maintained memory layer should evaluate Mem0, which provides a managed service and production API, or LangMem, which integrates with LangGraph. Both are included in LightMem's baseline evaluation framework, so benchmark comparisons are directly available.
LightMem's advantage over those alternatives for researchers is the explicit paper backing (ICLR 2026), the open implementation across multiple memory strategies (LightMem, StructMem, FluxMem, EM²Mem), and the standardized evaluation on LoCoMo and LongMemEval. The last push to the repository was on 2026-09-05. The MIT license permits unrestricted use and modification.
Editorial conclusion
LightMem is a good choice for researchers evaluating memory augmentation strategies and for developers building conversational agents that need persistent context beyond a single session. It is not suited for production deployments requiring strict version stability: the package is at version 0.1.0, Python is constrained to 3.10 or 3.11, and many dependencies are pinned to specific versions. Before adopting, run the reproduction scripts on LoCoMo and LongMemEval to verify that the baseline results hold on your hardware configuration.
Frequently asked questions
How do I add DeepSeek model support to LightMem?
The README notes that the configuration file at `src/lightmem/configs/memory_manager/base_config.py` accepts `deepseek-v4-flash` and `deepseek-v4-pro`, with `reasoning_effort` and thinking-mode configuration options added in April 2026.
What benchmarks does LightMem support for evaluation?
LightMem includes reproduction scripts for LoCoMo and LongMemEval at `experiments/locomo/readme.md` and `experiments/longmemeval/readme.md`. A baseline evaluation framework in `src/lightmem/memory_toolkits/readme.md` also supports running Mem0, A-MEM, and LangMem on both datasets for comparison.
Does LightMem include an MCP server for agent integration?
Yes. The MCP server at `mcp/server.py` exposes LightMem's memory tools to MCP-compatible agents. It was added in November 2025 and allows agents using the Model Context Protocol to read, write, update, and retrieve memories through standard tool calls.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/zjunlp-lightmem)