Model or dataset
Raudaschl/rag-fusion avatar
Raudaschl/rag-fusion

RAG-Fusion: Multi-Query Generation and Reciprocal Rank Fusion for Better Retrieval

RAG-Fusion: multi-query generation + Reciprocal Rank Fusion for better retrieval-augmented generation. Includes evaluation harness with NFCorpus/BEIR.

959 stars114 forksPythonMIT

At a glance

What is it?
RAG-Fusion is a Python repository that extends standard retrieval-augmented generation by generating multiple query rewrites with an LLM, running both BM25 and vector search across all of them, and merging the ranked lists with Reciprocal Rank Fusion before passing candidates to a reranker. An evaluation harness using NFCorpus and BEIR datasets measures the retrieval lift each configuration provides.
Who is it for?
RAG-Fusion is worth adopting when users phrase queries with vocabulary that differs from how the indexed corpus is written, when missing a document is more costly than surfacing a marginal one, and when a strong reranker is already part of the retrieval pipeline.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 3 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The Problem RAG-Fusion Addresses

Standard RAG relies on a single user query to retrieve context. When the user's phrasing does not match how the indexed documents are written, the embedding similarity between query and relevant passage is low. The relevant documents may still exist in the corpus, but they rank below less-relevant ones because their vocabulary diverges.

RAG-Fusion addresses this by generating several rewrites of the original query, each phrased from a different angle. These rewrites are then searched separately, and the resulting ranked lists are merged using Reciprocal Rank Fusion. Documents that appear in multiple lists accumulate score across appearances, so documents that are broadly relevant to the topic surface more reliably than documents that match only one phrasing.

The technique fits scenarios with terminology mismatch between user vocabulary and indexed text, academic or scientific literature where lay and technical language diverge, patent prior-art search, legal e-discovery, and exploratory research where the user does not know the correct terminology in advance. The README describes poor-fit cases explicitly: FAQ chatbots with curated content, voice or autocomplete systems with sub-second latency budgets, and code or identifier search where precision dominates recall.

How the Pipeline Works

The pipeline has four stages. First, an LLM generates four rewrites of the user's query, each approaching it from a different angle: synonyms, narrower framings, broader framings, and related sub-topics. Second, both BM25 and vector search run over the original query and each rewrite, producing multiple ranked lists. Third, Reciprocal Rank Fusion merges these lists. Each document in any list receives a score contribution of `1 / (60 + rank)`, and scores are summed across all appearances. The constant 60 dampens the effect of very high-ranked documents in single lists. Fourth, a reranker, either a cross-encoder or the Jev decision model, re-scores the top candidates from the fused pool before the answer model receives its context.

The original technique, demonstrated in `main.py`, uses only vector search without a reranker. The README states that experiments found this vector-only variant is roughly neutral once any reranker is added, and the version that reliably improves retrieval is the hybrid variant combining BM25 and vector search with a strong reranker.

Installing and Running the Evaluation

The repository requires an OpenAI API key for query rewriting, answer generation, and LLM-judge scoring. A `.env` file stores credentials:

bash
cp .env.example .env

Edit `.env` and replace `your-key-here` with the actual key. The `JEV_API_KEY` field in `.env.example` is only needed for the Jev reranker experiments.

Core dependencies for the basic demo and evaluation harness:

bash
pip install openai chromadb python-dotenv tqdm tabulate rank_bm25

Optional dependencies for reranking experiments:

bash
pip install sentence-transformers flashrank typesafe-sdk

The README notes that the LLM used for rewrites, answers, and judging defaults to `gpt-6-luna` at low reasoning effort, and can be overridden with the `LLM_MODEL` and `LLM_REASONING_EFFORT` environment variables. The evaluation harness runs on the NFCorpus dataset from the BEIR benchmark suite.

Repository Layout and the Experiments

The repository separates the original demonstration from the experimental evaluation. `main.py` implements the core pipeline as a standalone script. `evaluate.py` is the evaluation CLI for comparing retrieval configurations. `test_main.py` contains unit tests.

The `eval/` directory holds the components that support controlled experiments: `dataset.py` handles NFCorpus download and loading, `metrics.py` implements precision, recall, NDCG, and MRR metrics, `retrieval.py` covers BM25, vector, hybrid, and RAG-Fusion retrieval variants, `rerank.py` covers cross-encoders, FlashRank, and Jev, and `query_cache.py` persists rewrite results to disk to avoid re-running the LLM call on repeated runs.

The `experiments/` directory contains write-ups of two controlled studies. The September 2026 study tested four rerankers including Jev, documented cost and latency per configuration, and found a rewrite-parser bug and fixed it. The April 2026 study replicated an arXiv paper on RAG-Fusion and includes a correction note for the parser bug found later.

The README reports specific experimental findings: with `bge-reranker-large`, the hybrid diverse fusion variant showed `+0.025 NDCG@10` compared to baseline (n=200, 95% confidence intervals excluding zero). With Jev the lift was `+0.050 NDCG@10`.

Cost, Latency, and When Not to Use It

RAG-Fusion adds latency and a small cost to every query. The README estimates that the rewrite LLM call costs about $0.06 per 1,000 queries with a small model and adds approximately 1.5 seconds before the parallel searches can start. Adding BM25 alongside vector search, which is the hybrid baseline without rewrites, is described as close to free and worth doing regardless of whether fusion is used.

The README explicitly lists cases where RAG-Fusion is the wrong choice. FAQ chatbots and curated customer-support knowledge bases have precise document sets where the vocabulary mismatch problem does not occur. Latency-critical workloads such as voice assistants, autocomplete, or sub-second-p95 chat cannot absorb the rewrite delay. High-volume, margin-thin consumer search where the query LLM cost is not justified falls in the same category.

The README discusses query routing as a mitigation for mixed workloads: run fusion only on queries where it is likely to help, and skip it for queries where the vocabulary already matches the corpus. A routing experiment in the September 2026 study used Jev as a classifier, which fused 78% of queries and saved about a fifth of rewrite calls at no measurable retrieval cost, but did not improve over non-selective fusion because it did not target the queries where fusion actually helps. A router keyed on a retrieval-weakness signal such as a low top reranker score is described as the next thing to test.

Licence and Practical Scope

The repository is published under the MIT license. It is a reference implementation and evaluation framework, not a library intended for import. There is no PyPI package, no versioning scheme, and no documented API surface intended for downstream use. Teams that want to adopt the technique will read the code and adapt it to their own retrieval stack rather than adding a dependency on this repository.

The code in `main.py` demonstrates the original approach and does not include the hybrid search or reranking stages. The `eval/` directory is where the full pipeline lives. Researchers who want to replicate the experiments need the optional dependencies for reranking and access to both the OpenAI API and the NFCorpus dataset from BEIR.

Editorial conclusion

RAG-Fusion is worth adopting when users phrase queries with vocabulary that differs from how the indexed corpus is written, when missing a document is more costly than surfacing a marginal one, and when a strong reranker is already part of the retrieval pipeline. The repository's own experiments found that the hybrid variant combining BM25 and vector search with a reranker (specifically `bge-reranker-large` or Jev) produces the most reliable lift, while the vector-only variant is roughly a wash once any reranker is added. Teams with latency-critical requirements should note that the rewrite call adds approximately 1.5 seconds before the parallel searches can begin. The project is not a production library: it has no PyPI package, and the code in `main.py` demonstrates the original vector-only technique while the more effective hybrid configuration lives in `eval/`.

Frequently asked questions

What is RAG-Fusion?

RAG-Fusion is a retrieval technique that generates multiple rewrites of a user's query using an LLM, runs BM25 and vector search across all rewrites, and merges the ranked results with Reciprocal Rank Fusion. Documents that appear across multiple query rewrites accumulate score and surface higher than documents that match only one phrasing.

How does Reciprocal Rank Fusion work in this repository?

Each document in any ranked list receives a score of `1 / (60 + rank)`, where 60 is a constant that dampens the advantage of very high positions in a single list. Scores are summed across all lists the document appears in. The resulting fused ranking is then passed to a reranker before the final candidates reach the answer model.

Does the repository include a reranker out of the box?

The core demo in `main.py` does not use a reranker. The `eval/` directory supports four rerankers: cross-encoders via `sentence-transformers`, FlashRank, and the Jev decision model from TypeSafe. The optional dependencies `sentence-transformers`, `flashrank`, and `typesafe-sdk` are required for the reranking experiments.

Official sources

  1. Issues
  2. License: MIT
  3. Raudaschl/rag-fusion on GitHub
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/raudaschl-rag-fusion.svg)](https://hysenlabs.com/projects/raudaschl-rag-fusion)