open-rag-eval: RAG evaluation without golden answers
RAG evaluation without the need for "golden answers"
At a glance
- What is it?
- Vectara's Apache-2.0 Python toolkit scores RAG pipelines with UMBRELA and AutoNuggetizer, so you can run an evaluation before anyone has written reference answers. Here is what it does, how to install it, and where it stops.
- Who is it for?
- Adopt open-rag-eval if you already have a query set and want TREC-RAG style metrics without hand-writing golden answers, and if you are willing to pay for an LLM judge and, in the default configuration, hold a Hugging Face token with access to vectara/hallucination_evaluation_model.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 119 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem open-rag-eval targets: evaluation without a golden answer set
Most RAG evaluation workflows assume you have reference answers or labeled relevant chunks. Producing those is slow, and they go stale the moment you change your retriever or your prompt. open-rag-eval is built around the opposite premise. Its README states that the core metrics, UMBRELA and AutoNuggetizer, "do not require golden chunks or golden answers," which is what makes a first evaluation possible on a corpus where nobody has written labels yet. The techniques come from Jimmy Lin's lab at UWaterloo, and the toolkit implements the metrics used in the TREC-RAG benchmark.
The audience is narrow but real: engineers who own a RAG pipeline, already have a list of questions users actually ask, and want a score they can compare across two configurations. If you are still choosing a vector database, this is not the tool. If you have a query log and a pipeline that keeps changing, it is.
How the evaluation pipeline is wired: connectors, evaluators, artifacts
The architecture is three layers. A connector pulls answers out of your RAG system. An evaluator computes metrics over those answers. A results folder holds per-query scores and intermediate outputs so you can see why a query scored the way it did. The README describes the design as modular, with custom metrics and connectors as the extension points, and lists out-of-the-box connectors for the Vectara RAG platform, LlamaIndex and LangChain.
Configuration is a YAML file, and the example config exposes the seams directly: input_queries points at your CSV, generated_answers and eval_results_file name the intermediate and final outputs, results_folder collects the artifacts, and the connector block under options/query_config carries the corpus key and any Vectara query parameters. That structure matters because it means the evaluation is reproducible from two files, a query list and a config, rather than from a notebook someone edited by hand.
Two dependencies shape what you can actually run. The default LLM judge needs an OpenAI API key. Factual consistency, used for hallucination detection, has two paths: the open-source HHEM model, which is the default and needs a Hugging Face token plus access approval for vectara/hallucination_evaluation_model, or Vectara's commercial Factual Consistency API. That is a real fork in the road, not a footnote.
Installing open-rag-eval and running a first evaluation against Vectara
Install from pip. The README notes this also installs the open-rag-eval command-line tool, so you can run evaluations and plot results without cloning anything.
pip install open-rag-evalIf you want the example configs, or you intend to work on the code, build from source instead.
git clone https://github.com/vectara/open-rag-eval.git
cd open-rag-eval
pip install -e .Set the credentials the default metrics expect. The OpenAI key is required for the default judge, and the Hugging Face token is required if you take the default HHEM path for factual consistency.
export OPENAI_API_KEY='your-api-key'
export HF_TOKEN='your-huggingface-token'Write your queries into a CSV with a single column named query. The README gives this exact shape, and you can also have the toolkit generate queries for you.
query
What is a blackhole?
How big is the sun?
How many moons does jupiter have?Then edit config_examples/eval_config_vectara.yaml: point input_queries at your CSV, set generated_answers, eval_results_file and results_folder, and fill in your Vectara corpus_key in the connector section. The README says to customize any Vectara query parameter there to match the configuration you want to evaluate. What you should see afterwards is a results folder containing per-query scores and the intermediate outputs the reporting layer produces, which is where you go when a score looks wrong.
Where open-rag-eval gets awkward: judge cost, tokens and scoring drift
The no-golden-answer design does not remove the cost of evaluation, it moves it. UMBRELA and AutoNuggetizer are LLM-judged, so every run spends OpenAI tokens proportional to your query count, and the default factual consistency path pulls a transformer model locally. The requirements pin torch==2.7.1 and transformers==4.50.2, which means a first install is not a small download, and the Dockerfile is built on python:3.10-slim with build-essential added, so image builds are heavier than the runtime suggests.
The sharper limitation is comparability. Because the judge is a model, scores are only meaningfully comparable when the judge model, the prompt templates and the metric configuration are held fixed. The toolkit supports custom prompt templates, including one supplied by file path and one inlined in the YAML, and that flexibility cuts both ways: change the template between two runs and you have changed the measurement instrument, not the pipeline. Nothing in the README describes a frozen judge or a versioned scoring contract, so treat cross-release score comparisons as something you have to control yourself.
There is also a licensing boundary worth reading carefully. The open-source HHEM path requires requesting access to vectara/hallucination_evaluation_model on Hugging Face, and the alternative is a commercial Vectara API. So the default open path still depends on a gated model artifact, and the fully managed path depends on a paid account. Neither is a defect, but neither is the frictionless story the phrase "no golden answers" might suggest.
How it differs from Ragas and from plain retrieval metrics
Ragas is the obvious comparison for anyone evaluating a RAG pipeline in Python, and the difference is in what each treats as ground truth. Ragas-style workflows lean on reference answers or reference contexts for several of their metrics, which is exactly the labeling work open-rag-eval is designed to avoid. open-rag-eval instead derives its judgments from the query and the retrieved results through UMBRELA and AutoNuggetizer, with golden-answer evaluation available as an optional extra when references happen to exist. That makes open-rag-eval the better fit when you have queries but no labels, and Ragas the better fit when you already have a curated answer set you trust and want metrics anchored to it.
The second comparison is against retrieval-only evaluation. Recall@k, nDCG and MRR measure whether the right chunks came back, and they need relevance labels to mean anything. open-rag-eval is aimed at the end-to-end question of whether the generated answer is supported and useful, which is why factual consistency sits in the dependency list. If your problem is that the retriever returns the wrong chunks, a retrieval metric will localize that faster and cheaper than an LLM judge will.
Maintenance, releases and what adopting it commits you to
The repository is not archived, and the last push was on 2026-06-02. Releases have moved at a steady clip: v0.2.3 on 2025-10-27, v0.2.4 on 2025-11-18, and v0.3.0 on 2025-12-15. The setup.py classifier still reads "Development Status :: 4 - Beta", which is consistent with a toolkit whose extension points are still being filled in. The README says connectors for LlamaIndex and LangChain exist and that more are coming, so connector coverage is a moving target rather than a finished surface.
Upgrade cost is dominated by the dependency set, not by the package itself. requirements.txt pins torch, transformers, langchain, llama_index, pandas and pydantic at specific versions, so an upgrade of open-rag-eval can drag a large part of your environment with it. The Makefile gives the project's own checks: lint runs pylint and flake8, mypy runs type checking, and test runs unittest discovery under tests with TRANSFORMERS_VERBOSITY=error and TOKENIZERS_PARALLELISM=false. Both lint and mypy are suffixed with || true, so they report without failing the build, which is worth knowing before you treat a green run as proof of cleanliness.
The licence is Apache-2.0, which is permissive and includes an explicit patent grant. That covers the code in this repository. It does not automatically cover the gated Hugging Face model, the Vectara API, or the OpenAI judge you call, each of which carries its own terms. Read those separately rather than assuming the Apache-2.0 grant extends to them; this is an observation about where the boundaries sit, not legal advice.
Editorial conclusion
Adopt open-rag-eval if you already have a query set and want TREC-RAG style metrics without hand-writing golden answers, and if you are willing to pay for an LLM judge and, in the default configuration, hold a Hugging Face token with access to vectara/hallucination_evaluation_model. Skip it if you need a stable scoring contract across many releases, if you are evaluating a system whose retrieved chunks you cannot expose to the connector, or if you only want retrieval metrics such as recall and nDCG. Before committing, run the Vectara example end to end and inspect the per-query intermediate outputs, because those files are what tell you whether the judge is scoring your pipeline or its own assumptions.
Frequently asked questions
What are RAG evals?
RAG evals are measurements of how well a retrieval-augmented generation pipeline answers questions. open-rag-eval implements the metrics used in the TREC-RAG benchmark, including UMBRELA and AutoNuggetizer, and reports per-query scores plus intermediate outputs.
Is ChatGPT a RAG?
This question is not about open-rag-eval. The toolkit evaluates a RAG pipeline you already run, and its default LLM judge happens to call OpenAI through the openai package, but ChatGPT itself is not what the project measures.
What is RAG and why is it used?
The README does not explain RAG itself; it assumes you already run a retrieval-augmented generation pipeline and want to measure it. open-rag-eval focuses on evaluating the answers that pipeline produces, with optional factual consistency checking for hallucination detection.
What is RAG analysis and how does it work?
In open-rag-eval, analysis means scoring a set of queries against your RAG system and reviewing the resulting artifacts. A connector pulls answers from your pipeline, an evaluator computes metrics such as UMBRELA and AutoNuggetizer, and the results folder holds per-query scores and intermediate outputs for debugging.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/vectara-open-rag-eval)