What is Hallucination in LLMs?
Hallucination (also called AI hallucination) is when a large language model produces fluent, confident output that is not grounded in its training data, the prompt, or any real source. The model is not lying; it is filling gaps with plausible text, because its objective is to predict the next token, not to check facts.
How hallucination happens
A language model is a next-token predictor. Given a context, it assigns probabilities to possible continuations and samples from them. There is no separate step that verifies whether the sampled sentence corresponds to anything in the world. Fluency and factuality are produced by the same mechanism, so a sentence can be grammatical, on-topic and wrong at the same time.
Three conditions make this more likely. First, a gap in the context: if the prompt asks about something the model saw rarely or never during training, the probability mass still has to go somewhere, and the model produces the most plausible-sounding completion rather than an admission of ignorance. Second, a gap in the training data: long-tail entities, recent events, private codebases and internal APIs are underrepresented, so the model interpolates between similar patterns it has seen. Third, a decoding objective that rewards confidence: instruction-tuned models are trained to answer, and an answer that hedges or refuses is often scored worse than a fluent guess during preference tuning.
The failure is not random noise. It is structured. Models tend to invent citations that look like real citations, function names that follow the conventions of the library, and API arguments that match the shape of the surrounding code. That is exactly why hallucination is hard to spot by reading alone: the output is plausible in form and wrong in content.
When you need to care about it, and when you do not
Hallucination matters when the output is used as a fact or executed as a command. A generated citation, a legal clause, a medical dosage, a Terraform resource block or a shell command all carry a cost when wrong, and the cost is not proportional to how confident the text sounds. It also matters when the model summarizes a document, because a summary that adds a claim not present in the source is a hallucination even if the claim is true elsewhere.
It matters less when the output is a draft a human will verify against a known source, or when the task is open-ended generation where no ground truth exists. Brainstorming names, rewriting a paragraph, or producing variations on a theme do not require factual grounding. The distinction is not the model, it is the verification path: if a human or a tool checks the output against a source of truth before it is used, the cost of a hallucination is bounded.
The practical rule is to ask what happens if this specific sentence is wrong. If the answer is nothing, the risk is low. If the answer is a broken build, a wrong answer to a user, or a compliance problem, the output needs a grounding mechanism, not a better prompt.
Common pitfalls and limits
The first pitfall is treating a lower hallucination rate as solved. A model that hallucinates less often still hallucinates, and the failures cluster in the long tail, which is where users are most likely to be. A benchmark score is an average over a fixed test set; it does not bound the error on your data.
The second is confusing retrieval with grounding. Retrieving a document and pasting it into the prompt does not guarantee the answer comes from that document. The model can still blend retrieved text with parametric memory. Grounding requires either a citation back to a retrieved span or a check that the claim appears in the source.
The third is prompt-level suppression. Instructions like "do not hallucinate" or "only use the provided context" reduce the rate but do not remove the mechanism, because the model has no internal signal that distinguishes a remembered fact from a plausible completion. The same applies to self-critique: a model asked to check its own answer is using the same weights that produced the error.
The fourth is measurement scope. A leaderboard that tests summarization of short documents, such as vectara/hallucination-leaderboard, measures that task. Its README is explicit that it is only that. Reading it as a general hallucination ranking overstates what the test covers.
The fifth is silent failure. A hallucinated function call often fails loudly at runtime, but a hallucinated explanation fails quietly, and the reader has no signal to distinguish it from a correct one.
How it shows up in open-source projects
A group of open-source projects treats hallucination as a systems problem rather than a model problem, and their designs differ in where they put the check.
idosal/git-mcp turns a GitHub repository URL into a remote Model Context Protocol endpoint, so assistants read the project's own docs and code instead of guessing. The mechanism is context supply: the model is given the repository as a source rather than relying on parametric memory. The project's own description states it is not a code index and not a substitute for opening the source yourself, so it narrows the gap without closing it.
vectara/hallucination-leaderboard ranks LLMs by how often they invent content when summarizing short documents. It is a narrow, repeatable test, and the README is explicit that it is only that. It is useful for comparing models on one task and misleading if read as a general measure.
LukasNiessen/terrashark addresses a specific domain where the cost of a wrong answer is a broken plan. It is a MIT-licensed collection of markdown guidance that agents load before writing Terraform or OpenTofu, built around a small always-loaded SKILL.md plus 19 reference files read on demand. The design bet is that grounding in official HashiCorp best practices is cheaper than correcting invented resources and arguments after the fact.
Jane-xiaoer/claude-skill-web-clone routes website cloning through a decision tree by site type and grades uncertain WebGL reconstructions as SOURCE, PARTIAL or GUESS rather than presenting guesses with false confidence. The grading is the anti-hallucination mechanism: it makes uncertainty visible instead of hiding it behind fluent output.
beita6969/ScienceClaw layers 285 skills, persistent memory and a 629-line protocol onto the OpenClaw engine. Its notable rule forbids citing anything a tool did not return in the current session, which converts an unverifiable claim into a protocol violation.
MatrixOrigin/memoria adds snapshots, branches, merges and rollback on top of MatrixOne's copy-on-write engine. It targets agents that accumulate facts across sessions and need to undo a bad memory instead of retraining a prompt, which treats a stored hallucination as a state problem.
renee-jia/scholar-loop runs literature search, hypothesis generation, real PyTorch experiments, self-critique and write-up as one governed multi-agent loop. Its selling point is the deterministic harness around the agents, which is where the guards against reward-hacking and hallucination sit.
linghungegeg/Linghun is a local-first TypeScript CLI that puts model output behind evidence, permissions and verification. a5c-ai/babysitter is an MIT-licensed Node.js tool that runs the same enforced workflow across 12 AI coding harnesses and records every decision in an immutable journal. Johell1NS/browser-search wires SearXNG, Camofox and CloakBrowser into one escalation path so an agent can search and browse without API keys, with the README's strongest claims being the ones worth checking hardest.
What actually reduces it
No single technique removes hallucination. The projects above suggest a layered approach: supply the source (git-mcp, TerraShark), make uncertainty explicit (claude-skill-web-clone, ScienceClaw), verify against a tool or test (scholar-loop, Linghun, babysitter), and make state reversible (memoria). Each layer catches a different failure. Context supply fails when the source is incomplete; grading fails when the grader is wrong; verification fails when the checker shares the model's blind spot; rollback fails when the bad memory has already been acted on.
The constraint is cost. Every verification step adds latency and complexity, and a check that is wrong in the same direction as the model adds confidence without adding safety. The useful question is not which model hallucinates least, but where the output enters a decision and what verifies it before that happens.
In practice
Hallucination is a property of next-token prediction, not a bug that a better prompt removes. Treat model output as a draft until something outside the model confirms it: a retrieved source, a test, a tool result or a human review. To go further, read the README of vectara/hallucination-leaderboard to see what a narrow measurement looks like, and idosal/git-mcp to see how grounding a model in a repository's own files changes what it can answer.