# vectara/hallucination-leaderboard: what the HHEM numbers actually measure

> Vectara's public leaderboard ranks LLMs by how often they invent content when summarizing short documents. It is a narrow, repeatable test, and the README is explicit that it is only that.

**vectara/hallucination-leaderboard** — Leaderboard Comparing LLM Performance at Producing Hallucinations when Summarizing Short Documents

- Repository: https://github.com/vectara/hallucination-leaderboard
- Website: https://vectara.com
- Stars: 3,312 · Forks: 106
- Language: Python
- License: Apache-2.0
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/vectara-hallucination-leaderboard

## The problem: summarization models that add facts the source never contained

A summarization model that paraphrases badly is annoying. A summarization model that states a number, a name or a causal claim that is absent from the source document is a different class of failure, because the output reads exactly like a correct summary. In a retrieval-augmented pipeline the summary is often the only thing a downstream reader or system sees, so an invented detail propagates without any visible signal.

The hallucination leaderboard exists to make that failure countable. It evaluates how often an LLM introduces hallucinations when summarizing a document, using Vectara's Hallucination Evaluation Model, referred to as HHEM. The audience is narrow and practical: engineers picking a summarization model, and anyone who needs a published, dated number rather than a vendor claim. The README describes the project as a public LLM leaderboard computed with HHEM, and states an intent to update it as the model and the LLMs change. That intent is visible in the data: the table is stamped "Last updated on May 11, 2026", and the version history is kept in branches rather than overwritten.

## How the leaderboard is built: HHEM scoring over a fixed summarization task

The mechanism is a fixed task plus a scoring model. Each LLM is asked to summarize short documents, and HHEM then judges the summaries for factual consistency against the source. The published table exposes four columns per model: Hallucination Rate, Factual Consistency Rate, Answer Rate, and Average Summary Length in words.

The arithmetic is transparent. Hallucination Rate and Factual Consistency Rate sum to 100 percent in every row, so they are two views of one measurement. Answer Rate is the fraction of prompts where the model produced a summary at all, and it is the column most readers skip. It should not be skipped. microsoft/Phi-4 shows a 3.7 percent hallucination rate next to an 80.7 percent answer rate, and snowflake/snowflake-arctic-instruct shows 4.3 percent next to 62.7 percent. A model that declines to summarize most of the time scores well on the metric that only counts the summaries it did produce. Average Summary Length matters for the same reason: the table spans 54.7 words for openai/gpt-5.4-mini-2026-03-17 to 254.4 for openai/gpt-5.1-high-2025-11-13, and a longer summary has more surface area on which to be judged inconsistent.

The repository itself is small. The top level holds CITATION.cff, LICENSE, README.md and an img directory, and the primary language is Python. The results table lives inline in the README between HTML comment markers, and a plot of the top 25 hallucination rates is committed as a PNG under img/. There is no service to run and no API to call. The README points readers to an interactive version of the leaderboard on Hugging Face, and to two earlier branches: hhem-1.0-final for the first version based on HHEM-1.0, and hhem-2.3-old-dataset for the version based on the previous dataset.

## Reading the table without misreading it

The table is sorted by hallucination rate, which makes the top rows look like a quality ranking. They are a ranking on one task. A model at 1.8 percent and a model at 3.1 percent are separated by a difference that the README gives no confidence interval for, and the model identifiers are dated snapshots such as openai/gpt-5.4-nano-2026-03-17 and anthropic/claude-sonnet-4-20250514. A provider can ship a new snapshot under a similar name and the number moves.

Two practical habits follow from the column set. First, read Answer Rate before Hallucination Rate, because a low hallucination rate paired with a low answer rate describes a cautious model, not an accurate one. Second, read Average Summary Length alongside both, because summarization verbosity is a style parameter as much as a capability. The README does not document per-model prompt settings or whether length was constrained, so a reader cannot tell from the table alone how much of a length difference is the model's default behaviour and how much is configuration. That is a real gap in the published material, and it limits how far the ranking can be pushed.

## Getting the data and citing it

There is no package to install and no CLI. The README does not give install steps because the project publishes results, not software you run. The two things you actually consume are the table in the README and the interactive leaderboard on Hugging Face.

The results table is delimited in the README by HTML comment markers, LEADERBOARD_START and its closing counterpart. That marker pair is the anchor any script should target when parsing the table out of the file.

If you cite the leaderboard in internal documentation, the repository ships a CITATION.cff at the top level for that purpose. There is no release artifact: the releases list is empty, so the branch head is the only version you can pin. The CITATION.cff file is the authoritative source for how the maintainers want the work attributed.

For a first real use, open the interactive leaderboard on Hugging Face, find the model you are considering, and read its Answer Rate and Average Summary Length before you read its rank. Those two columns change how much weight the headline hallucination rate deserves.

## Where the leaderboard stops being the right tool

The scope is summarization of short documents. Nothing in the README claims the metric transfers to long-context synthesis, multi-document question answering, tool use, code generation, or open-ended chat. A model that scores well here can still fabricate a citation in a chat response, and the leaderboard would not show it.

The scoring model is also a single judge. HHEM is Vectara's own evaluation model, and the leaderboard is computed by the same organization that publishes it. That is not disqualifying, and the version history is unusually transparent for a leaderboard, but it does mean the metric is one model's judgement of factual consistency, not a ground-truth annotation by independent raters. The README does not describe an inter-annotator agreement study or a human audit of HHEM's verdicts.

The third limit is the update cadence. The last push to the repository was on 2026-05-11, and the table carries the same date. Between updates, a model you are evaluating may have been superseded by a newer snapshot that is not in the table at all. Treating the ranking as current-model guidance months after the stamp is a misuse of the artifact.

## Alternatives and how they differ in approach

The closest alternative in kind is a general-purpose academic hallucination benchmark such as the HaluEval line of work, which asks models to detect or generate hallucinated content across several task types rather than scoring one summarization task with one judge. The difference is breadth against comparability: HaluEval-style suites cover more failure modes, while this leaderboard keeps the task fixed so that a new model's number is directly comparable to the previous one.

A second alternative is to skip public leaderboards and run your own evaluation on your own documents with a judge you control. That is more work and it is not comparable across teams, but it removes both the single-judge problem and the short-document scope limit. If your production summaries are of 40-page contracts rather than short documents, this is the option that actually answers your question.

A third is to use HHEM directly as a runtime guardrail rather than as a comparison table. The leaderboard and the scoring model are separate artifacts; the leaderboard is a snapshot of model behaviour, while HHEM is the thing that produced it. The README does not document how to deploy HHEM in a pipeline, so that route requires going to the model's own documentation rather than this repository.

## Maintenance, licence and what an upgrade costs you

The repository is not archived, and the last push was on 2026-05-11. That is the whole of the maintenance signal available here: there are no tagged releases, so there is no version to upgrade between and no changelog to read. "Upgrading" means pulling the branch and diffing the README table against the copy you parsed last time. If you have built a parser around the LEADERBOARD_START marker, that is a small job. If you have hard-coded model names from a previous read, expect churn, because the table lists dated snapshots and providers rename or retire them.

The licence is Apache-2.0, which is permissive and includes an explicit patent grant. The repository also ships a CITATION.cff, which signals that the maintainers expect academic citation rather than silent reuse. Apache-2.0 does not require attribution in the way CC BY does, but the presence of a citation file is a clear statement of intent. This is not legal advice, and if you are republishing the table or the plot at scale you should read the LICENSE file and the CITATION.cff yourself rather than relying on a summary.

## Conclusion

Use this leaderboard if you are choosing a summarization model for a retrieval pipeline and you want a consistent, published number to compare candidates before you run your own evaluation. Do not use it as a general measure of model truthfulness, and do not treat the ranking as stable across model versions: the table lists dated model snapshots, and a new snapshot can move a model several points. Before adopting a number, check the Answer Rate column, confirm the model identifier matches the exact version you would deploy, and read the hhem-1.0-final and hhem-2.3-old-dataset branches if you need to understand how the metric changed between dataset versions.

## FAQ

### Does AI still hallucinate in 2026?

According to the leaderboard table dated May 11, 2026, every listed model has a non-zero hallucination rate, ranging from 1.8 percent for antgroup/finix_s1_32b upward. The table does not include any model at 0.0 percent.

### What AI hallucinates the most on the Vectara hallucination leaderboard?

The table is sorted by hallucination rate, so the highest rates appear at the bottom. Among the rows shown, xai-org/grok-4-fast-non-reasoning sits at 19.7 percent and mistralai/ministral-3-14b-2512 at 19.4 percent.

### What is the hallucination rate of LLMs according to the Vectara hallucination leaderboard?

Rates in the May 11, 2026 table span from 1.8 percent to 19.7 percent across the models listed, with Factual Consistency Rate as the complement. The leaderboard measures hallucination when summarizing short documents, not across all tasks.

### What is the hallucination rate on the AA Omniscience test?

The repository does not mention the AA Omniscience test, so the leaderboard cannot answer this. The only metric it publishes is HHEM-based hallucination rate for document summarization.

## Sources

- [Issues](https://github.com/vectara/hallucination-leaderboard/issues)
- [License: Apache-2.0](https://github.com/vectara/hallucination-leaderboard/blob/main/LICENSE)
- [Project website](https://vectara.com)
- [README](https://github.com/vectara/hallucination-leaderboard/blob/main/README.md)
- [vectara/hallucination-leaderboard on GitHub](https://github.com/vectara/hallucination-leaderboard)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/vectara-hallucination-leaderboard
