Model or dataset
hhy-huang/HiRAG avatar
hhy-huang/HiRAG

HiRAG: Hierarchical Knowledge Graph RAG for Multi-Hop Questions

[EMNLP'25 findings] An easy-to-use Graph RAG system using hierarchical knowledge.

557 stars83 forksPythonMIT

At a glance

What is it?
A Python RAG library, accepted to EMNLP 2025 Findings, that organises a text corpus into a three-level knowledge hierarchy and retrieves from all levels in a single query. It outperforms LightRAG and GraphRAG on the UltraDomain benchmark across four domains, but re-indexing is slow and the dependency list is heavy.
Who is it for?
HiRAG is a good fit for teams building retrieval systems over large, domain-specific corpora such as legal documents, research papers, or agricultural knowledge bases, where questions require connecting facts from multiple parts of the text. Teams that need quick re-indexing when the source corpus changes will find the upfront graph construction cost a meaningful bottleneck, since the README explicitly flags re-indexing cost as a motivation for a companion project.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 105 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What HiRAG Adds to Standard RAG

Standard chunked RAG splits a corpus into fixed-size segments, embeds each one, and retrieves the nearest chunks to a query. That works well for factual lookups but breaks down on questions that require synthesising information from different parts of a document or across documents. Multi-hop questions, broad thematic queries, and requests for comprehensive answers across a knowledge base all suffer from the flat retrieval model.

HiRAG addresses this by building a knowledge graph during indexing and organising extracted knowledge at three levels. Local knowledge captures entity-level facts. Global knowledge captures community-level summaries derived from clusters of related entities. The hi retrieval mode queries all three levels and merges the results before passing them to the language model. This design is what the authors describe as hierarchical, and it is the central claim of the EMNLP 2025 paper.

The Three-Level Indexing and Retrieval Architecture

During the insert step, HiRAG sends the text to a language model to extract named entities and their relationships. Those entities form a knowledge graph stored in a neo4j instance. A community detection algorithm from the graspologic library then clusters the graph and generates community-level summaries using the LLM, creating the global layer. Low-level chunks serve as the local layer.

At query time, the hi mode retrieves candidates from all three layers, combines them, and passes the merged context to the generation model. The README also exposes four narrower modes for comparison: naive (plain chunk retrieval), hi_nobridge (hierarchical without the bridge between local and global), hi_local (hierarchical local only), and hi_global (hierarchical global only). These modes exist to support ablation experiments.

The working directory stores all intermediate state so that LLM calls can be cached across runs with enable_llm_cache=True. The README notes that re-indexing a knowledge base is expensive, which is why the authors also released a companion project called DeepRefine for test-time knowledge base refinement.

Installing and Running a First Query

The repository must be cloned before installing because the package is not published to PyPI separately. Clone and install in editable mode:

bash
cd HiRAG
pip install -e .

The quick-start example in the README shows the minimal setup:

python
graph_func = HiRAG(
    working_dir="./your_work_dir",
    enable_llm_cache=True,
    enable_hierachical_mode=True,
    embedding_batch_num=6,
    embedding_func_max_async=8,
    enable_naive_rag=True
    )

After instantiation, insert a text file and run a query:

python
with open("path_to_your_context.txt", "r") as f:
    graph_func.insert(f.read())
print(graph_func.query("The question you want to ask?", param=QueryParam(mode="hi")))

For DeepSeek, the repository provides hi_Search_deepseek.py as a ready-made entry point. The corresponding files for OpenAI and ChatGLM are hi_Search_openai.py and hi_Search_glm.py. Cohere and Ollama variants are also present in the repository. API keys and model names go in config.yaml at the repository root.

Running the Evaluation Pipeline

The eval/ directory contains a four-step pipeline used in the paper. The README documents the Mix dataset as the example:

shell
cd ./HiRAG/eval
python extract_context.py -i ./datasets/mix -o ./datasets/mix

After extracting context, insert it into the graph database. Then run queries with a specific mode:

shell
python test_deepseek.py -d mix -m hi

The batch_eval.py script handles LLM-based pairwise evaluation in two phases: request (sending comparisons) and result (collecting scores). On the Mix dataset, the paper reports HiRAG at 87.6% overall against Naive RAG at 12.4%, 64.1% against GraphRAG at 35.9%, and 65.9% against LightRAG at 34.1%, using pairwise comparison by a judge LLM. These numbers come directly from the README tables. The evaluation methodology uses an LLM judge rather than reference labels, which means the scores reflect relative preference rather than an objective ground truth.

Limitations and When to Use a Different System

The dependency list in requirements.txt is substantial and version-pinned: graspologic 3.4.1, neo4j 5.25.0, openai 1.61.1, umap_learn 0.5.6, transformers 4.47.1, and others. Installing these in an environment that already has incompatible versions of numpy or scikit-learn will produce conflicts. The pinned numpy 1.26.4 in particular conflicts with environments that have moved to NumPy 2.

Re-indexing is the largest operational cost. Every time the source corpus changes substantially, you rebuild the graph and regenerate community summaries. The README flags this explicitly. For corpora that update frequently, this is a poor fit. The companion DeepRefine project is described as addressing test-time refinement, but HiRAG itself does not support incremental graph updates.

The tool also requires a running neo4j instance. For teams running containerised workloads this is manageable, but it rules out quick experimentation on a laptop without Docker.

How HiRAG Differs from LightRAG

LightRAG is a graph-based RAG system that also extracts entities and relationships from text and builds a knowledge graph for retrieval. Both projects target multi-hop and complex question answering over long documents. The difference is in how the graph is queried and how community-level knowledge is handled.

HiRAG explicitly organises retrieved knowledge into a three-level hierarchy and queries all levels in a single pass. LightRAG offers local and global query modes but treats them as separate paths rather than as a unified hierarchical structure. The benchmark tables in the HiRAG README show that on the Mix, CS, Legal, and Agriculture datasets from UltraDomain, HiRAG wins the pairwise comparison against LightRAG in every dimension and every dataset. The gaps are widest on the Mix and CS datasets and narrowest on Legal and Agriculture.

GraphRAG from Microsoft is a third alternative. It also builds community summaries from a knowledge graph. The HiRAG README reports that HiRAG outperforms GraphRAG on all four datasets in the benchmark, though the margin is smaller on the Legal and Agriculture domains than on Mix and CS.

Editorial conclusion

HiRAG is a good fit for teams building retrieval systems over large, domain-specific corpora such as legal documents, research papers, or agricultural knowledge bases, where questions require connecting facts from multiple parts of the text. Teams that need quick re-indexing when the source corpus changes will find the upfront graph construction cost a meaningful bottleneck, since the README explicitly flags re-indexing cost as a motivation for a companion project. Verify that your infrastructure can run neo4j and graspologic before adopting it: both add operational overhead that naive chunked RAG avoids. The last commit to the repository was on 2026-06-16.

Frequently asked questions

What is hierarchical RAG and how does it work?

Hierarchical RAG organises a corpus into multiple knowledge levels during indexing and retrieves from all levels at query time. HiRAG builds a three-level structure: entity-level local knowledge, community-level global summaries derived from graph clustering, and the raw text chunks. The hi retrieval mode queries all three and merges the results before generation.

What LLM backends does HiRAG support?

The repository includes ready-made entry-point files for DeepSeek (hi_Search_deepseek.py), OpenAI (hi_Search_openai.py), ChatGLM (hi_Search_glm.py), Cohere (hi_Search_cohere.py), and Ollama (hi_Search_ollama.py). API keys and model names are set in config.yaml.

What datasets was HiRAG evaluated on?

The evaluation uses four domains from the UltraDomain dataset available on Hugging Face: Mix, CS (computer science), Legal, and Agriculture. The eval/ directory includes scripts for extracting, inserting, testing, and scoring against all four domains.

Official sources

  1. hhy-huang/HiRAG on GitHub
  2. Issues
  3. License: MIT
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/hhy-huang-hirag.svg)](https://hysenlabs.com/projects/hhy-huang-hirag)