Hallucination & Grounding
Why models confidently produce false statements, and the operational patterns that reduce the failure rate.
Preparing spaced-repetition scheduler…
A hallucination is a model output that is fluent or confident but not supported by the provided context, reliable sources, or verifiable evidence. There are three common kinds:
- Closed-domain hallucination — the model is given a document and asked to answer from it, but invents a fact that contradicts or goes beyond the document. One mitigation is grounding: instruct the model to answer only from the provided context, and to say "I don't know" otherwise.
- Open-domain hallucination — the model is asked a factual question with no source. The model may produce a plausible-but-wrong answer, especially for niche or recent topics beyond its training cutoff. RAG and citation checks can reduce this risk: provide relevant sources and verify that they support the answer. Citations alone are not proof.
- Reasoning hallucination — the model produces internally inconsistent or arithmetically wrong steps, but presents the final answer as correct. Mitigate it with structured decomposition and external verifiers, such as a calculator for math or executable tests for code.
Why does it happen? The model is a next-token predictor, not a fact database. When its training data is sparse on a topic, or when the prompt pushes it into an unfamiliar region of the input space, the most probable next token is not the most true one. Confident, fluent text is a property of the language modeling objective — it does not imply confidence in truth.
Operational mitigations:
- RAG with citations for fresh or proprietary facts.
- Tool use for verifiable computation (search, code execution, calculators).
- Self-consistency: compare multiple answers, then verify externally; correlated errors can win a majority vote.
- Refusal training: teach the model to say "I don't know" when context is insufficient.
- Evaluation: build a labeled eval set and track hallucination rate over time.
Key takeaways
- Hallucination is fluent-but-ungrounded output from a next-token predictor.
- RAG, tool use, and refusal training are the main mitigations.
- Always build an evaluation set and measure hallucination rate over time.
Self-check
What is a closed-domain hallucination?
Why is self-consistency (sample-many, take-majority) a partial mitigation?
References
- Why language models hallucinate
Research on factual errors, uncertainty, and evaluation incentives.
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
Original RAG research; retrieval and parameter updates can be combined.
- Model optimization
Evaluate prompting, retrieval, and fine-tuning against the task.
Last verified 2026-09-04.