LEANN: a 97% smaller vector index for local RAG
[MLsys2026 Best Paper]: https://arxiv.org/abs/2506.08276. RAG on Everything with LEANN. Enjoy 97% storage savings while running a fast, accurate, and 100% private RAG application on your personal device.
At a glance
- What is it?
- LEANN is a Python vector database that recomputes embeddings on demand instead of storing them, so a personal RAG index can fit on a laptop. The trade-off is that recomputation happens at query time, and the README does not document what that costs.
- Who is it for?
- Adopt LEANN when the index has to live on the same machine as the data and you cannot ship embeddings to a hosted store: personal file systems, mail archives, browser history, agent memory. Do not adopt it when you need a managed service with replication, or when you cannot tolerate embedding computation inside the query path, since the README describes recomputation as the mechanism and does not publish a latency budget.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 25 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The storage problem LEANN is aimed at
A conventional vector database stores one float vector per chunk. The README gives the arithmetic for a large corpus: 60 million text chunks occupy 201GB in a traditional setup and 6GB with LEANN. That ratio is where the 97% figure comes from, and it is the whole reason the project exists. The audience is not a company running a shared retrieval service. It is someone with 200GB of PDFs, mail, and chat logs on a laptop, who wants semantic search over all of it without renting a GPU box or uploading the corpus to a vendor. The README frames this as "RAG on Everything" and lists file systems, Apple Mail, browser history, WeChat and iMessage history, ChatGPT and Claude conversation archives, Slack, Twitter bookmarks, and codebases as targets. The constraint that ties those together is that the data is personal and the machine is the same one it lives on. LEANN is MIT licensed and the repository is not archived; the last push was on 2026-09-05.
Selective recomputation and high-degree preserving pruning
The README names graph-based selective recomputation with high-degree preserving pruning as the mechanism. Instead of persisting an embedding for every chunk, LEANN keeps a proximity graph and computes embeddings on demand while traversing it. The pruning step is what keeps that graph small: it preserves high-degree nodes, which are the hubs a graph search relies on to move between clusters, and drops the rest. Storage therefore shifts from dense float arrays to a sparse graph plus CSR-format adjacency, which the README cites as the reason memory use stays low as well. The cost model is inverted relative to FAISS-style indexes. You pay at query time, not at ingest time. That is a defensible design for a personal device, where disk is scarce and queries are occasional, and a bad one for a service answering thousands of queries per second. The README does not publish a latency comparison against a stored-embedding index, so the query-time cost is the number you have to measure yourself.
Installing LEANN and indexing a folder
The README requires uv first and gives the install script for it. After that, the documented path is to clone the repository for the examples and install the package from PyPI. Note that the install command in the README is truncated mid-name ("uv pip install lea"), so copy the package name from the PyPI page rather than from the README snippet.
curl -LsSf https://astral.sh/uv/install.sh | sh
git clone https://github.com/yichuan-w/LEANN.git leann
cd leann
uv venv
source .venv/bin/activate
uv pip install leannThe repository ships runnable examples rather than a single quickstart. examples/basic_demo.py is the entry point for a first index, and examples/dynamic_update_no_recompute.py is the one to read if your corpus grows, since its name states the property being demonstrated: updates without a full recompute. For a code corpus, examples/grep_search_example.py and examples/llamaindex_hybrid_example.py cover the keyword and hybrid paths, and benchmarks/contextbench/README.md is where the project says the Claude Code comparison can be reproduced. Run the examples from the cloned directory, since they import from the repository layout, and check the README in the repository for the exact invocation for each script.
Expect the first run to download an embedding model, because the dependencies include sentence-transformers and transformers. That download happens once and is cached locally, which is consistent with the offline-first claim, but it does mean the first index build is not offline.
What the storage claim does not cover
The 97% number is a storage comparison, and the README is explicit that it is one. Nothing in the README states a query latency figure for LEANN against a stored-embedding baseline, and the ContextBench numbers are about coding-agent retrieval quality, not about speed. The agent result is a separate claim: on 30 SWE-Bench Pro tasks, with model, agent, tools, and an 8,192-token retrieval budget held fixed, the project reports 2.1x initial relevant-code recall (24.2% against 11.4% for BM25), 12.6 percentage points higher coverage after exploration (38.4% against 25.8%), and 8.4% fewer tokens (3.22M against 3.51M). The README attaches its own caveat to that: better context access does not guarantee issue resolution. Treat the recall gain as a retrieval-quality result and the storage figure as the only efficiency claim on offer.
The second gap is operational. Embeddings computed on demand mean the embedding model has to be resident and fast enough that traversal does not stall. On an ARM64 Mac the dependency list pulls mlx and mlx-lm; on other platforms you are relying on torch and sentence-transformers. If your embedding model is large, the recomputation path is where the time goes, and the README does not tell you how much.
When a stored-embedding index is the better choice
FAISS is the honest comparison, and it is already in the dependency tree through llama-index-vector-stores-faiss. FAISS stores the vectors and searches them directly. That makes ingest expensive in disk terms and query cheap and predictable. LEANN does the opposite: ingest is cheap in disk terms and query carries the embedding work. The choice follows from which resource you are short of. A workstation with a fast NVMe drive and a service-level latency target should keep the vectors and use FAISS. A laptop with 512GB of storage and one user issuing a few queries a minute is the case LEANN was built for. There is also a hybrid path in the repository, examples/llamaindex_hybrid_example.py, which suggests the project expects some users to combine LEANN with a conventional index rather than replace it outright. If your corpus is small enough that a full embedding matrix fits comfortably, LEANN's graph machinery buys you nothing and adds a recomputation step.
Maintenance, upgrades, and the MIT licence
The repository is not archived and the last push was on 2026-09-05, so the project is being touched. Releases are infrequent rather than continuous: v0.3.5 on 2025-11-12, v0.3.6 on 2026-01-11, v0.3.7 on 2026-03-08. That cadence matters if you pin versions, because there is no patch stream between releases and a bug you hit may sit until the next minor version. The workspace pyproject.toml pins protobuf to exactly 4.25.3 and constrains transformers to >=4.53.1,<4.58 on Python 3.10 and above, with a separate <4.46 pin for Python 3.9. Those pins are the upgrade cost: moving transformers forward means checking the tree-sitter and astchunk code-chunking dependencies that sit alongside it. Python 3.10 through 3.14 are the supported interpreters.
The MIT licence covers the code. It does not cover the embedding models you point it at, and it does not cover the data you index. Those carry their own terms, and the README's privacy claim is about where computation happens, not about what you are permitted to index.
Editorial conclusion
Adopt LEANN when the index has to live on the same machine as the data and you cannot ship embeddings to a hosted store: personal file systems, mail archives, browser history, agent memory. Do not adopt it when you need a managed service with replication, or when you cannot tolerate embedding computation inside the query path, since the README describes recomputation as the mechanism and does not publish a latency budget. Before committing, verify three things on your own corpus: the actual index size against a FAISS baseline, query latency with your embedding model on your hardware, and whether the backend you pick (leann-backend-hnsw is the one declared in pyproject.toml) supports the update pattern you need. The 97% figure is a storage claim from the project, not a latency claim.
Frequently asked questions
What is leanness?
In this article the term refers to LEANN's storage property: the index keeps a pruned proximity graph and recomputes embeddings on demand rather than storing one vector per chunk, which the README quantifies as 6GB against 201GB for 60 million chunks.
How do I install LEANN?
Install uv first, then clone the repository and create a virtual environment with uv venv, activate it, and install the package from PyPI. The README's install line is truncated, so take the package name from the PyPI page.
Does LEANN work offline?
The repository is tagged offline-first and the README states that data never leaves the laptop. The first index build still downloads an embedding model, because sentence-transformers and transformers are dependencies.
Which Python versions does LEANN support?
The README badge lists Python 3.10, 3.11, 3.12, 3.13, and 3.14, and pyproject.toml requires Python 3.10 or newer with a separate transformers pin for 3.9 compatibility.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/startrail-org-leann)