Model or dataset
DocAILab/XRAG avatar
DocAILab/XRAG

XRAG: benchmarking the core modules of a RAG pipeline, one component at a time

XRAG: eXamining the Core - Benchmarking Foundational Component Modules in Advanced Retrieval-Augmented Generation

729 stars20 forksPythonApache-2.0

At a glance

What is it?
XRAG is a Python framework that swaps retrievers, embeddings, LLMs and evaluation metrics in and out of a RAG pipeline so you can measure which module actually changes the answer. It ships as the examinationrag package with an xrag-cli entry point and a Streamlit WebUI.
Who is it for?
Adopt XRAG if you need to attribute a RAG quality change to one module rather than to the whole pipeline, and you are willing to run the pinned dependency set on Python 3.9 or newer. Do not adopt it if you want a production retrieval service or a thin library to import into an existing app; the console script and Streamlit UI are the intended surfaces.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What XRAG is for: isolating one module at a time

Most RAG evaluation reports a single end-to-end score. That number moves when you change the splitter, the embedding model, the retriever, the reranker or the prompt, and it does not tell you which of them moved it. XRAG's stated purpose is to dissect those modules and measure how each configuration affects overall performance, which is a different job from shipping a retrieval service. The audience is implied by the packaging: the classifiers in setup.py list Developers, Education and Science/Research, and the repository carries an arXiv badge (2412.15529) and a note that the accompanying paper was accepted at ICDE 2026. This is a measurement harness for people writing about or tuning RAG systems, not a runtime dependency for an application. The unit of work is a comparison: same dataset, same questions, one module changed, metrics reported per configuration.

The pluggable pipeline: retrievers, post-processors and orchestrators

The README describes a modular design with pluggable components for retrievers, embeddings and LLMs. On the retrieval side the documented options are vector search, BM25, hybrid and tree-structured retrieval, plus keyword-based retrieval and document summary retrieval. Post-processing is a separate stage, and requirements.txt pins three rerankers for it: Cohere rerank, ColBERT rerank and a FlagEmbedding reranker. Text splitting was added later (SemanticSplitterNodeParser, SentenceSplitterNodeParser, SentenceWindowNodeParser, per the 2025-07-18 update), which matters because chunking is one of the modules most often blamed for retrieval failures. Orchestrators sit above all of this and control execution order. The README names five kinds: sequential, conditional, iterative, parallel and hybrid, and the update log records self-rag and adaptive-rag arriving on 2025-11-05 and SIM-rag on 2025-11-18. Evaluation is likewise layered: traditional metrics (F1, EM, MRR, Hit@K, MAP, NDCG), LLM-based metrics (Faithfulness, Relevancy, Correctness) and what the README calls deep evaluation (Contextual Precision/Recall, Hallucination, Bias), drawing on LlamaIndex, DeepEval and custom metrics. The practical consequence is that a comparison run produces several metric families at once, and disagreement between them is itself information.

Installing examinationrag and running a first evaluation

The distribution name on PyPI is examinationrag, while the import package and console script are xrag and xrag-cli. setup.py declares python_requires >=3.9.0 and defines the entry point as xrag-cli = xrag.cli:main, so the CLI is the primary surface. The README does not give a pip install line; the package is published on PyPI under the name examinationrag, so installation goes through the distribution name rather than the import name. There is an optional extra for the jury dependency group, declared in setup.py as extras_require = {'jury': ['jury']}. For a first real use, the repository ships examples/data/ and examples/generated_qa.json, and the 2025-01-05 update added a generate command that produces QA pairs from a folder of documents. That is the entry point to start from, because every later metric needs questions with reference answers. Interactive evaluation runs through the WebUI, which the README launches with this command:

bash
xrag-cli webui

The README's workflow for that interface is five steps: upload and configure a dataset (it names HotpotQA, DropQA and NaturalQA, and says custom datasets go through automatic format conversion), configure the API key and build the vector index including chunk size, define the pipeline (pre-retrieval methods, retriever, post-processor, prompt template), test queries interactively with retrieval results and generated responses visible, then run the full evaluation. Configuration lives in config.toml at the repository root, with a default_config.toml bundled inside the package as package data; the 2025-06-23 update reworked that file so parameters read more clearly. The 2025-01-09 update added an API mode, so the same components can be driven as a backend service rather than only through the CLI or UI.

Where XRAG gets in the way

The dependency set is the first obstacle. requirements.txt pins llama-index==0.11.11 and a matching set of llama-index subpackages (embeddings-openai 0.2.5, llms-openai 0.2.9, retrievers-bm25 0.3.0, readers-file 0.2.2, postprocessor-colbert-rerank 0.2.1), alongside torch==2.3.0, transformers==4.42.2, sentence_transformers==3.0.1 and deepeval==0.21.26. Those are exact pins, not ranges. If your application already depends on a different llama-index release, installing examinationrag into the same environment is a conflict, and a virtual environment per comparison run is the sane default. Second, this is an evaluation harness, not a serving layer. There is no documented path for using XRAG's retrievers inside a latency-sensitive query path, and the WebUI is a Streamlit app, which is a reasonable interface for a human running experiments and a poor one for concurrent production traffic. Third, the roadmap is honest about what is unfinished: the checklist still has Semantic Perplexity, Entropy and Semantic Entropy, Auto-J, Prometheus, late chunking, FLARE and multi-step approaches open, and the LLM experiment list notes 30B-or-larger models via OpenAI API or Ollama as a pending item. If your comparison depends on one of those, the framework does not yet provide it. Fourth, the README does not document rollback, checkpointing or resuming a partially completed evaluation, so a long run that fails has no stated recovery path. Finally, the last release is v0.1.4 from 2025-02-07 while the repository's last push was 2026-06-03, so the published package lags the source tree; features described in the update log after February 2025 (text splitters, self-rag, adaptive-rag, SIM-rag) are not in the PyPI release.

XRAG against a general-purpose evaluation library

The nearest comparison is DeepEval, which requirements.txt already pins at 0.21.26. The difference is scope. DeepEval is a metric library: you bring a pipeline, it scores the outputs. XRAG owns the pipeline as well, so the thing being varied is inside the framework. That is why the retrievers, splitters, rerankers and orchestrators are all first-class, and why the WebUI walks through index construction and strategy configuration rather than just metric selection. The cost of that design is coupling: to compare two retrievers in XRAG you work inside XRAG's configuration model, whereas with a metric library you would change one line in your own code and rerun. A second comparison point is RAGAS-style metric suites, which also assume you supply the pipeline. XRAG's answer to that is the orchestrator layer, which lets a run be sequential, conditional, iterative, parallel or hybrid, matching the agentic RAG patterns the project is aiming at. If your question is "is my system's faithfulness acceptable", a metric library is lighter. If your question is "did the reranker or the splitter cause this", XRAG's structure is the point.

Licence, maintenance and upgrade cost

The project is Apache-2.0, declared in setup.py as "Apache 2.0 License" and classified as an OSI-approved Apache Software License. Apache-2.0 permits commercial use and modification and includes a patent grant; it also requires that you retain the licence and notice files and state significant changes. That is a summary of the licence text, not legal advice, and the arXiv paper or any benchmark data you publish from a run may carry its own terms that the repository licence does not cover. On maintenance: the repository is not archived and the last push was 2026-06-03, which is recent relative to the project's own history. The release cadence is a different story. v0.1.2, v0.1.3 and v0.1.4 all landed within about a month in January and February 2025, and nothing has been released since. The update log shows continued work in the source tree through November 2025 and a paper acceptance note dated 2026-02-24, so the gap is between the published package and the repository, not between the repository and inactivity. Upgrade cost is dominated by the exact pins: moving to a newer examinationrag means moving llama-index, torch and transformers together, and any comparison you ran under the old pin set is not directly comparable under the new one unless you rerun the baseline.

Editorial conclusion

Adopt XRAG if you need to attribute a RAG quality change to one module rather than to the whole pipeline, and you are willing to run the pinned dependency set on Python 3.9 or newer. Do not adopt it if you want a production retrieval service or a thin library to import into an existing app; the console script and Streamlit UI are the intended surfaces. Before committing, verify that your dataset converts into the QA format the generate command produces, that your chosen LLM backend (OpenAI, Ollama, DashScope, ZhipuAI or a local HuggingFace model) is reachable from your environment, and that the pinned llama-index 0.11.11 line does not conflict with the version your application already uses.

Frequently asked questions

What is XRAG and what does it evaluate?

XRAG is a benchmarking framework for the foundational component modules of retrieval-augmented generation systems, packaged as examinationrag. It evaluates retrieval quality, response faithfulness and answer correctness using traditional metrics such as F1, EM, MRR, Hit@K, MAP and NDCG, LLM-based metrics such as Faithfulness, Relevancy and Correctness, and deep evaluation metrics including Contextual Precision/Recall, Hallucination and Bias.

How do I install XRAG and launch its WebUI?

Install the PyPI distribution named examinationrag, then run the console script xrag-cli webui to start the Streamlit interface. setup.py declares the entry point as xrag-cli = xrag.cli:main and requires Python 3.9 or newer.

Which LLMs and retrievers does XRAG support?

The README lists OpenAI models and local models such as Qwen and LLaMA, and requirements.txt pins clients for DashScope, ZhipuAI, Ollama and HuggingFace. Retrieval options documented in the README are vector, BM25, hybrid, tree-based, keyword-based and document summary retrieval, with optional Cohere, ColBERT and FlagEmbedding rerankers as post-processors.

Does XRAG generate its own QA pairs from my documents?

Yes. The 2025-01-05 update added a generate command that produces QA pairs from a folder containing your documents, and the repository ships examples/data/ and examples/generated_qa.json as reference output.

Can XRAG be used as a backend service instead of through the CLI?

The 2025-01-09 update added API support, described in the README as using XRAG as a backend service. requirements.txt includes fastapi, uvicorn and pydantic, which is consistent with that mode.

Official sources

  1. DocAILab/XRAG on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/docailab-xrag.svg)](https://hysenlabs.com/projects/docailab-xrag)