# LightMem: A Lightweight Memory Framework for LLMs and AI Agents

> LightMem is a Python framework from the zjunlp group, accepted at ICLR 2026, that adds long-term memory to LLM applications and AI agents. It provides a storage, retrieval, and update layer that works with both cloud APIs and local models, and ships with an MCP server for agent integration.

**zjunlp/LightMem** — [ICLR 2026] LightMem: Lightweight and Efficient Memory-Augmented Generation

- Repository: https://github.com/zjunlp/LightMem
- Stars: 1,182 · Forks: 114
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/zjunlp-lightmem

## The Long-Context Memory Problem LightMem Addresses

LLM agents lose context between sessions because they have no mechanism to store and retrieve information from past interactions. Within a single session, the context window eventually fills, and relevant earlier content is dropped. LightMem addresses this by providing an explicit memory layer: a modular system for storing conversation history, retrieving relevant memory entries for each new query, and updating or expiring memories over time.

The primary intended use is building AI applications that require long-term memory: conversational assistants that remember user preferences and facts across sessions, research agents that accumulate findings over multiple runs, and multi-agent systems where agents need to share or access a common memory store. The README describes the framework as targeting both LLM developers building intelligent applications and AI safety researchers benchmarking memory strategies on standard datasets.

## Memory Storage, Retrieval, and Update Architecture

LightMem's architecture separates memory into three operations: storage, retrieval, and update. Storage handles writing new memories derived from conversation turns. Retrieval selects relevant memories for a given query, using a combination of retrieval strategies that can include BM25 keyword matching, vector similarity (via `qdrant-client`), and LLM-based compression with `llmlingua`.

The framework supports a modular storage backend design: the README describes custom storage engines and retrieval strategies as supported through the modular architecture. The pyproject.toml lists `qdrant-client==1.15.1` as a dependency, indicating vector-based memory retrieval is the primary backend for semantic search.

The LongMemEval evaluation specifically tests offline memory update, meaning the system updates stored memories based on conversation history before answering questions. This mode is documented in the `experiments/longmemeval/readme.md` reproduction script and is distinct from online update, where memories are updated incrementally during a live conversation session.

The repository also includes StructMem, a hierarchical memory variant that preserves event-level memory bindings and cross-event connections (accepted at ACL 2026), FluxMem, a connectivity-evolving memory framework that models memory as a heterogeneous graph (under review at EMNLP 2026), and EM²Mem, an event-centric multimodal memory framework accepted at EMNLP 2026. Each is documented in its own markdown file in the repository root.

## Installing LightMem and Running the Baseline Evaluation

LightMem requires Python 3.10 or 3.11 (the pyproject.toml specifies `>=3.10,<3.12`). The package is available on PyPI.

```bash
pip install lightmem
```

For adding DeepSeek model support, the README notes the configuration file at `src/lightmem/configs/memory_manager/base_config.py`, which accepts `deepseek-v4-flash` and `deepseek-v4-pro` with `reasoning_effort` and thinking-mode parameters (added April 2026).

For local model support via Ollama or vLLM, the factory implementations are at `src/lightmem/factory/memory_manager/ollama.py` and `src/lightmem/factory/memory_manager/vllm_offline.py`. A Transformers auto-loading path is at `src/lightmem/factory/memory_manager/transformers.py`.

The baseline evaluation framework at `src/lightmem/memory_toolkits/readme.md` supports running Mem0, A-MEM, and LangMem on LoCoMo and LongMemEval for comparison. Reproduction scripts for LightMem's own results on LoCoMo are at `experiments/locomo/readme.md` and for LongMemEval at `experiments/longmemeval/readme.md`. Tutorial notebooks in `tutorial-notebooks/` cover common usage scenarios.

## MCP Server and Agent Integration

LightMem ships an MCP server at `mcp/server.py`, added in November 2025. This makes LightMem's memory tools available to MCP-compatible agents, including coding agents that support the Model Context Protocol.

The MCP server exposes LightMem's operations (read, write, update, retrieve) as tools that an agent can call during its execution. This integration path is relevant for teams building agents with Claude Code, OpenClaw, or other MCP-compatible orchestrators that need persistent memory without managing a separate memory service.

The repository also includes a `web/` directory, indicating a web interface for interacting with the memory system. The tutorial notebooks in `tutorial-notebooks/` cover multiple usage scenarios and were released alongside a demo video in December 2025. The CCF ODTC open source incentive program selected LightMem in July 2026, which indicates the project has received recognition from the Chinese open-source community beyond its ICLR acceptance.

## Benchmarks: LoCoMo and LongMemEval

LightMem includes evaluation results on two long-term memory benchmarks. LoCoMo tests memory retrieval over long social media conversation threads. LongMemEval tests both retrieval and offline memory update on long question-answering tasks.

The repository provides scripts for reproducing both sets of results, along with a unified baseline framework that runs Mem0, A-MEM, and LangMem on the same datasets for comparison. The README describes leading results on LoCoMo with strong performance and efficiency, directing readers to the reproduction scripts for the specific numbers.

The zjunlp group maintains a separate companion repository, MemBase, which provides a more comprehensive baseline evaluation framework supporting additional memory systems (including EverMemOS) on both LoCoMo and LongMemEval. This companion repository was announced in March 2026 and extends the comparison surface beyond what is included in the LightMem repository itself.

The dependency pin strategy (torch==2.8.0, transformers==4.57.0, specific library versions throughout) reflects the typical research artifact approach of locking dependencies for reproducibility rather than maintaining forward compatibility.

## Scope Boundaries and Alternatives

LightMem is a research framework. Version 0.1.0 and the tight Python version constraint (3.10 or 3.11 only) signal that the project prioritizes reproducibility and research flexibility over production deployment stability. Teams building production agents that need a supported, maintained memory layer should evaluate Mem0, which provides a managed service and production API, or LangMem, which integrates with LangGraph. Both are included in LightMem's baseline evaluation framework, so benchmark comparisons are directly available.

LightMem's advantage over those alternatives for researchers is the explicit paper backing (ICLR 2026), the open implementation across multiple memory strategies (LightMem, StructMem, FluxMem, EM²Mem), and the standardized evaluation on LoCoMo and LongMemEval. The last push to the repository was on 2026-09-05. The MIT license permits unrestricted use and modification.

## Conclusion

LightMem is a good choice for researchers evaluating memory augmentation strategies and for developers building conversational agents that need persistent context beyond a single session. It is not suited for production deployments requiring strict version stability: the package is at version 0.1.0, Python is constrained to 3.10 or 3.11, and many dependencies are pinned to specific versions. Before adopting, run the reproduction scripts on LoCoMo and LongMemEval to verify that the baseline results hold on your hardware configuration.

## FAQ

### How do I add DeepSeek model support to LightMem?

The README notes that the configuration file at `src/lightmem/configs/memory_manager/base_config.py` accepts `deepseek-v4-flash` and `deepseek-v4-pro`, with `reasoning_effort` and thinking-mode configuration options added in April 2026.

### What benchmarks does LightMem support for evaluation?

LightMem includes reproduction scripts for LoCoMo and LongMemEval at `experiments/locomo/readme.md` and `experiments/longmemeval/readme.md`. A baseline evaluation framework in `src/lightmem/memory_toolkits/readme.md` also supports running Mem0, A-MEM, and LangMem on both datasets for comparison.

### Does LightMem include an MCP server for agent integration?

Yes. The MCP server at `mcp/server.py` exposes LightMem's memory tools to MCP-compatible agents. It was added in November 2025 and allows agents using the Model Context Protocol to read, write, update, and retrieve memories through standard tool calls.

## Sources

- [Issues](https://github.com/zjunlp/LightMem/issues)
- [License: MIT](https://github.com/zjunlp/LightMem/blob/main/LICENSE)
- [README](https://github.com/zjunlp/LightMem/blob/main/README.md)
- [zjunlp/LightMem on GitHub](https://github.com/zjunlp/LightMem)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/zjunlp-lightmem
