LightRAG: a graph-based RAG server you can install with uv
[EMNLP2025] "LightRAG: Simple and Fast Retrieval-Augmented Generation.
At a glance
- What is it?
- LightRAG builds a knowledge graph from your documents and serves retrieval over HTTP. It is a Python package, ships a WebUI, and asks you to bring your own LLM and embedding keys.
- Who is it for?
- Adopt LightRAG if you want a self-hosted retrieval server where the knowledge graph is the point, and you are willing to run an LLM and an embedding model behind it. Skip it if you need plain vector search over a few documents, since the extraction pass costs model calls that a flat index does not.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 4 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
What LightRAG does that a vector store does not
A conventional retrieval pipeline splits documents into chunks, embeds them, and returns the nearest chunks to a query. LightRAG keeps that index but adds a second one: a knowledge graph of entities and relations extracted from the same text by an LLM. Queries can then be answered from entity neighbourhoods as well as from chunk similarity, which is the difference the paper linked from the repository homepage is about.
The intended user is a developer who already has an LLM endpoint and a document collection, and wants retrieval as a service rather than as a library call inside one script. The repository ships a server, a WebUI under lightrag_webui/, a Dockerfile, and a docker-compose.yml that publishes port 9621. That is a deployment shape, not a notebook demo. The cost of the graph is real: every inserted document triggers extraction calls to your model, so ingestion is slower and more expensive than embedding alone, and the quality of the graph depends on the model's ability to emit structured output.
How indexing and querying actually flow
The README's news list and the repository layout describe four storage roles that LightRAG fills separately: a key-value store, a vector store, a graph store, and a document status store. The graph store can be networkx (the default dependency), Neo4j, or OpenSearch, which the 2026.03 entry describes as a unified backend covering all four roles. PostgreSQL and MongoDB are also listed as all-in-one options in earlier entries.
Insertion goes through a chunking step. Four strategies are selectable as of 2026.05: Fix, Recursive, Vector, and Paragraph. Chunks go to the extraction prompt, which returns entities and relations; those land in the graph store, while embeddings land in the vector store. Query time merges both signals, and a reranker is supported and set as the default query mode according to the 2025.08 entry. Four LLM roles can be configured independently: EXTRACT, QUERY, KEYWORDS, and VLM. That split matters in practice, because extraction is the token-hungry role and you may want a cheaper model there than for final answer synthesis.
Installing LightRAG and running a first query
The README recommends uv over pip and gives the install line for the server extra. Run it, then create the environment file the server reads at startup.
uv tool install "lightrag-hku[api]"
cp env.example .envThe env.example file is not in the package; the README says to download it from the repository root or copy it from a local checkout. Open .env and set your LLM and embedding configuration before starting anything. The README does not enumerate every key in this section, so read the file itself.
For a container deployment, the repository provides a compose file that mounts ./data/rag_storage, ./data/inputs and ./data/prompts, and maps the host port through a PORT variable defaulting to 9621.
docker compose up -dWith the stack up, the WebUI is served from the same port, and documents dropped into data/inputs become ingestible. The examples directory covers the library path if you prefer not to run the server: examples/lightrag_openai_demo.py, examples/lightrag_ollama_demo.py and examples/lightrag_gemini_demo.py each wire a different provider, and examples/insert_custom_kg.py shows inserting a graph you built yourself rather than one the extractor produced. Before exposing the port, set LIGHTRAG_API_KEY or AUTH_ACCOUNTS with TOKEN_SECRET in .env. The docker-compose.yml comment states that without it every endpoint is public.
Where LightRAG is the wrong choice
The extraction step is the failure surface. If your documents are short, homogeneous, or already structured, the graph adds latency and model cost without adding recall, and a plain embedding index will answer the same questions. If your LLM is weak at producing valid JSON, extraction degrades, and the json_repair dependency in pyproject.toml exists precisely because malformed model output is expected often enough to need repair.
The project also labels itself Development Status :: 4 - Beta in pyproject.toml, and the release history shows a run of release candidates (v1.5.7rc1 and v1.5.7rc2) in the days before v1.5.7rc2 on 2026-08-19. Pinning is therefore not optional. The README does not document a rollback procedure for a failed graph rebuild, and the document deletion feature added in 2025.08 triggers automatic KG regeneration, which means a delete can be as expensive as an insert. Plan for that in your ingestion budget. Python 3.10 is the floor in pyproject.toml, and the pandas constraint there is deliberately widened to 4.0.0 so 3.10 resolves to 2.x while newer interpreters get 3.x.
LightRAG compared with GraphRAG
The comparison people search for is LightRAG vs GraphRAG, and the README's own framing is speed and simplicity. Microsoft's GraphRAG builds community summaries over the entity graph and answers global questions by summarizing those communities; that produces strong thematic answers at the cost of an expensive indexing pass. LightRAG's retrieval side, as described in the repository, combines graph neighbourhoods with vector hits and a reranker rather than relying on pre-computed community summaries.
The practical difference is where the cost sits. GraphRAG front-loads spend into indexing. LightRAG spreads it across extraction at insert time and retrieval at query time, and lets you assign a cheaper model to the EXTRACT role than to QUERY. If your queries are mostly local and entity-centric, that split is an advantage. If you need reproducible global summaries over a static corpus, the community-summary approach is the more direct fit. The repository also lists sibling projects from the same group, MiniRAG and RAG-Anything, and the 2026.05 entry says RAG-Anything was merged in for multimodal parsing through MinerU or Docling services.
Licence, upgrades and what maintenance looks like
The licence is MIT, declared in both LICENSE and the license field of pyproject.toml. That permits commercial use and modification with attribution, but it also means no warranty and no support obligation from the authors. If you embed LightRAG in a product, the MIT notice must travel with it; anything beyond that is a question for your own counsel, not for this article.
Upgrade cost is the part to watch. The last push to the default branch was on 2026-08-19, and the three most recent releases are v1.5.6 on 2026-08-06 followed by v1.5.7rc1 and v1.5.7rc2. A project shipping release candidates that close together is moving, and the storage layer has churned: OpenSearch arrived in 2026.03, role-specific LLM config and chunking strategies in 2026.05, smart heading recognition in 2026.07. Each of those touches configuration keys. Pin lightrag-hku to an exact version, keep your .env under version control separately from the package, and re-read env.example on every bump, because a new role key or storage option that is absent from your .env will fall back to a default you did not choose.
Editorial conclusion
Adopt LightRAG if you want a self-hosted retrieval server where the knowledge graph is the point, and you are willing to run an LLM and an embedding model behind it. Skip it if you need plain vector search over a few documents, since the extraction pass costs model calls that a flat index does not. Before committing, verify three things on your own corpus: that your LLM produces valid graph extraction JSON at the temperature you configure, that the storage backend you chose is the one listed in the storage matrix, and that LIGHTRAG_API_KEY is set in .env before you publish the port, because the docker-compose.yml comment states every endpoint is public without it.
Frequently asked questions
What is LightRAG?
It is a retrieval-augmented generation system from HKUDS that indexes documents into both a vector store and a knowledge graph, then serves queries over HTTP through a built-in server and WebUI. The repository describes it as simple and fast RAG, and links an EMNLP2025 paper on the homepage.
How do I install LightRAG?
The README recommends uv and gives the command uv tool install "lightrag-hku[api]" for the server, followed by copying env.example to .env and filling in your LLM and embedding settings. A pip path is documented as an alternative, and a Dockerfile plus docker-compose.yml are in the repository root.
What are the key differences between LightRAG and GraphRAG?
GraphRAG answers global questions through pre-computed community summaries, which concentrates cost in indexing. LightRAG, as the repository describes it, combines graph neighbourhood retrieval with vector search and a reranker, and lets you assign separate models to the EXTRACT and QUERY roles.
Is LightRAG open source and free?
Yes. The licence is MIT, declared in the LICENSE file and in pyproject.toml, so the code can be used and modified commercially with attribution. The models you connect it to are billed separately by their providers.
Is LightRAG slow?
Ingestion is slower than plain embedding because each chunk goes through an LLM extraction pass to build the graph, and document deletion triggers automatic KG regeneration. The 2025.10 entry describes work to remove processing bottlenecks for large datasets, and a reranker is the default query mode.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/hkuds-lightrag)
Community notes