AI Education

LLM Hallucinations: Why They Happen and How to Reduce Them

Illustration: a robot at a laptop is asked when the Eiffel Tower was built, hallucinating a fabricated 1886 date on one side while a verified source confirms the correct 1889 date on the other

Hallucination is the term for when a language model states something false with the same confidence it uses for something true. OpenAI’s own research team put it concretely: they asked a widely used chatbot for the title of a researcher’s PhD dissertation and got three different confident answers, none correct; asked for his birthday, they got three different dates, all wrong. Their conclusion, published in a September 2025 paper, isn’t that this is a mysterious flaw — it’s that standard training and evaluation reward confident guessing over honest uncertainty, and until that changes, hallucination doesn’t go away just because a model gets bigger.

This guide is built on that paper, OpenAI’s own GPT-5 system card benchmark data, and Vectara’s independently maintained hallucination leaderboard — not on vague claims about “AI making things up.”

Two different things people call “hallucination”

Worth separating before anything else, because the fix is different for each:

  1. Factual (closed-book) hallucination — a model states something false about the world from memory alone, with no source text to check against. “What’s this person’s birthday?” answered confidently and wrong is this kind.
  2. Faithfulness hallucination — a model is given a source document (in a summary, or a RAG pipeline) and states something that contradicts or isn’t supported by that source. This is what RAG is specifically designed to reduce, and it’s measured differently — against the provided text, not against the world.

The rest of this guide focuses mainly on the first kind, since it’s the one OpenAI’s research directly explains the mechanism for — but the benchmark section covers both.

Why guessing beats honesty on a leaderboard

OpenAI’s core argument is a scoring-incentive problem, and it’s easiest to see through their own multiple-choice analogy: if you don’t know an answer but guess anyway, you might get lucky. Leaving it blank guarantees zero. A model asked someone’s birthday and forced to guess “September 10” has roughly a 1-in-365 chance of being right; saying “I don’t know” guarantees a score of zero. Averaged across thousands of test questions, the model that always guesses ends up looking better on a leaderboard than the model that honestly abstains — even though it’s wrong far more often.

OpenAI’s own GPT-5 system card makes this concrete with real numbers from the SimpleQA benchmark:

Infographic: OpenAI's SimpleQA benchmark comparing gpt-5-thinking-mini against OpenAI o4-mini — the older o4-mini has a marginally higher accuracy rate of 24 percent versus 22 percent, but a far higher error rate of 75 percent versus 26 percent, because it abstains on only 1 percent of questions versus 52 percent
Metric gpt-5-thinking-mini OpenAI o4-mini
Abstention rate (no answer given) 52% 1%
Accuracy (correct answer) 22% 24%
Error rate (wrong answer — hallucination) 26% 75%

Read the raw accuracy column alone, and the older o4-mini looks slightly better — 24% versus 22%. Read the error rate, and it’s the opposite story: o4-mini hallucinates nearly three times as often, because it almost never says “I don’t know.” Strategic guessing under uncertainty improves accuracy while making hallucination worse — and most leaderboards report only accuracy.

OpenAI’s proposed fix, stated directly in the paper: score confident wrong answers more harshly than admitted uncertainty, and give partial credit for appropriate abstention — the same negative-marking idea some standardized tests have used for decades to discourage blind guessing. Their point isn’t that a good hallucination-specific eval doesn’t exist (it does) — it’s that as long as the mainstream, accuracy-only leaderboards keep rewarding guesses, models will keep learning to guess.

Where hallucination actually comes from

The deeper question: why does this specific type of error happen, when large models rarely make spelling mistakes or mismatched parentheses anymore?

The answer is in what pretraining data does and doesn’t label. A language model learns by predicting the next word across enormous amounts of text — and unlike a labeled dataset, no sentence comes tagged “true” or “false.” The model only ever sees positive examples of fluent language and has to approximate the overall distribution it came from.

OpenAI’s own analogy makes this concrete: image classifiers trained on millions of photos labeled “cat” or “dog” learn to classify reliably, because the pattern is learnable. But if you instead labeled every pet photo with that pet’s birthday, no amount of algorithmic sophistication would help — birthdays are arbitrary, not a learnable visual pattern. Spelling and matched parentheses behave like the cat/dog case: the pattern is consistent, so errors vanish as models scale up. Arbitrary, low-frequency facts — a specific person’s birthday, a specific paper’s exact title — behave like the birthday-from-photo case: no amount of scale makes them predictable from patterns alone, and that’s where hallucination concentrates.

Five claims about hallucination the research pushes back on

The paper closes by directly addressing common assumptions, and it’s worth walking through because a few of them are genuinely counterintuitive:

  • “Better accuracy will eventually eliminate hallucination.” No — accuracy will never reach 100% no matter how large, well-searched, or good at reasoning a model gets, because some real-world questions are genuinely unanswerable with the available information.
  • “Hallucination is unavoidable.” Also no — a model can abstain when it’s uncertain. The problem isn’t capability, it’s incentive.
  • “Avoiding hallucination needs a smarter, bigger model.” Not necessarily — a smaller model can find it easier to recognize its own limits. Asked something in a language it never trained on, a model with zero exposure can flatly say “I don’t know”; a model with partial exposure has to correctly judge its own confidence, which is a harder problem. Being well-calibrated about your own uncertainty takes much less computation than being consistently accurate.
  • “Hallucination is a mysterious failure mode.” It isn’t — the statistical mechanism (unlabeled pretraining data plus accuracy-maximizing evaluation) is understood well enough to predict where it concentrates.
  • “We just need a good hallucination benchmark.” Good ones already exist. They don’t move the needle much on their own, because they sit alongside hundreds of traditional accuracy-based evaluations that still reward guessing. The fix has to reach the mainstream evals, not just add a specialized one next to them.

The independent number: Vectara’s hallucination leaderboard

OpenAI’s paper explains closed-book factual hallucination. For the faithfulness kind — does a model’s summary of a given document stay accurate to that document — the most-cited ongoing benchmark is Vectara’s Hughes Hallucination Evaluation Model (HHEM) leaderboard, which scores models on how often they introduce unsupported claims while summarizing short documents, and has been updated continuously since its first release.

Infographic: Vectara's hallucination leaderboard for document summarization, showing hallucination rates from under 2 percent for the best-performing models up to around 10 percent for several well-known frontier models, as of September 2026

As of the leaderboard’s 22 September 2026 update, hallucination rates on this specific, narrower task (staying faithful to a provided document) range from under 2% for the best-scoring models up to roughly 10% for several well-known current frontier models — a useful reminder that faithfulness hallucination and closed-book factual hallucination are measured completely differently, and a model doing well on one isn’t guaranteed to do well on the other.

What actually reduces hallucination today

None of these eliminate it — nothing does, per OpenAI’s own first myth above — but each addresses a specific part of the problem:

  • Retrieval-augmented generation — grounding answers in retrieved source text turns a closed-book factual question into an open-book one, which is the faithfulness problem HHEM measures, not the harder unanswerable-question problem. See our RAG guide for how this works end to end.
  • Structured outputs — doesn’t reduce factual hallucination, but eliminates a related failure mode: a schema-valid response can still contain a wrong value, but it can’t contain a malformed one. See our Structured Outputs guide for where this helps and where it doesn’t.
  • Reasoning models — OpenAI states plainly that GPT-5 hallucinates “significantly less often, especially while reasoning,” though not never. Explicit reasoning steps give a model more opportunity to check its own intermediate claims before committing to a final answer.
  • Prompting for abstention explicitly — telling a model it’s acceptable, even preferred, to say “I don’t know” or ask a clarifying question measurably shifts behavior, because the model’s training already contains examples of calibrated uncertainty — it’s the evaluation incentive during training and fine-tuning that currently suppresses it.
  • Citations and verifiable sources — asking a model to cite where a claim comes from doesn’t stop hallucination by itself (a citation can be fabricated too), but it converts an unverifiable claim into a checkable one for a human or a second automated pass.

Frequently asked questions

Why do LLMs hallucinate instead of just saying “I don’t know”?

Because they’re trained and evaluated on procedures that score a confident wrong guess the same as, or better than, an honest “I don’t know” — per OpenAI’s own research, most accuracy-based benchmarks give zero credit for abstaining and a chance at credit for guessing, so guessing wins on average.

Will hallucination go away as models get bigger?

Not on its own. OpenAI’s research states accuracy will never reach 100%, since some questions are inherently unanswerable, and scale doesn’t fix the incentive structure that rewards guessing over calibrated uncertainty.

Does RAG eliminate hallucination?

No — it reduces one specific kind, by giving the model a source document to ground its answer in, which is the “faithfulness” hallucination measured by benchmarks like Vectara’s HHEM. It doesn’t fix closed-book factual hallucination on questions with no retrieved source at all.

Are hallucination rates the same across all tasks?

No. A model’s hallucination rate on document summarization (Vectara’s HHEM measures this) and its hallucination rate on open-ended factual questions (OpenAI’s SimpleQA-style evals measure this) are different tasks, scored differently, and don’t necessarily correlate.

What’s the single biggest fix OpenAI proposes?

Changing how mainstream accuracy-based leaderboards score answers — penalizing confident wrong answers more than admitted uncertainty, and giving partial credit for appropriate abstention — rather than treating hallucination as something only a dedicated hallucination benchmark should measure.