LLM hallucination detection means automatically flagging answers that are unsupported by the source material or simply false. When you have a reference, such as retrieved documents in a RAG system, the most reliable approach is a faithfulness check that tests each claim against that context, using an LLM judge or a dedicated model like Vectara's HHEM. Without a reference, you can estimate risk by sampling several answers and measuring how much they disagree, using methods such as SelfCheckGPT or semantic entropy. No method catches everything, so production systems combine checks with human review.

This guide explains the main detection methods, when each one works, the open-source tools that implement them, and the benchmarks used to measure hallucination.

What counts as a hallucination?

It helps to separate two failure types, because they need different detectors.

  • Unfaithful answers (intrinsic hallucinations): the model contradicts or goes beyond the context it was given, for example a summary that adds a figure not in the document, or a RAG answer that cites a policy the retrieved pages never mention.
  • Factual errors (extrinsic hallucinations): the model states something false about the world with no source to check against, such as an invented citation, a wrong date or a non-existent API.

Faithfulness is far easier to check automatically, because you have the evidence in hand. Open-world factuality needs either a trusted knowledge source to look things up or an indirect signal such as the model's own uncertainty.

Hallucination detection methods compared

MethodNeeds a reference?How it worksMain weakness
Claim-level faithfulness with an LLM judgeYesSplit the answer into claims and check each against the contextJudge cost and judge errors
NLI or specialist detector modelsYesA trained model scores whether the context supports the answerWeaker outside its training domain
Self-consistency samplingNoSample several answers; disagreement signals hallucinationConsistent errors go unnoticed
Semantic entropyNoCluster sampled answers by meaning; high entropy flags confabulationSeveral generations per question
Retrieval-based fact checkingExternal sourceLook claims up in search or a knowledge baseOnly as good as the source
Human reviewEitherExperts verify sampled outputsSlow and expensive

To summarize: if your application has source documents, use faithfulness checks; if it answers from the model's own knowledge, use sampling-based uncertainty or external fact checking; and use human review to calibrate whatever automated method you choose.

Faithfulness checks with an LLM judge

The workhorse method breaks an answer into individual statements and asks a judge model whether each is supported by the context. The faithfulness score is the share of supported statements. Ragas implements this as its faithfulness metric (Ragas docs), DeepEval offers both a faithfulness metric for RAG and a hallucination metric that compares output with provided context (DeepEval docs), and TruLens calls it groundedness, one leg of its RAG triad. Because these rely on judges, calibrate them against human labels as described in our LLM-as-a-judge guide. For the full set of retrieval metrics, see our RAG evaluation metrics guide.

Specialist detector models

Instead of prompting a general model, you can use a model trained specifically to judge whether a text is supported by a source. Vectara's Hallucination Evaluation Model, known as HHEM, is openly available on Hugging Face and scores factual consistency between a source and a summary (model card). Patronus AI released Lynx, an open hallucination evaluation model (paper), and Patronus AI offers judge models as a hosted service. Specialist models are cheaper and faster than large judges at volume, but check them on your own domain before trusting them.

Self-consistency and semantic entropy

When there is no reference, one signal is whether the model agrees with itself. SelfCheckGPT samples several responses to the same prompt and checks whether the claims in the main answer are supported by the other samples; claims that appear in only one sample are likely invented (paper). Semantic entropy, published in Nature in 2024, refines this by clustering sampled answers by meaning rather than wording, then measuring uncertainty across the clusters; high semantic entropy flags confabulations (Nature). Open-source packages such as CVS Health's UQLM bundle uncertainty-based scorers like these.

The catch is that a model can be confidently and consistently wrong, and every check costs several extra generations.

Hallucination benchmarks and leaderboards

Benchmarks measure how often models hallucinate in a controlled setting:

  • TruthfulQA tests whether models repeat common human misconceptions (paper).
  • HaluEval is a large collection of generated and human-annotated hallucinated samples for testing whether models can recognize hallucinations (paper).
  • The Vectara Hallucination Leaderboard uses HHEM to measure how often models introduce unsupported content when summarizing documents, and reports hallucination rate, factual consistency rate and answer rate. It was last updated in September 2026.

These are useful for choosing a model, but a model's benchmark hallucination rate will not match your application's, which depends on your prompts, retrieval quality and domain. Our LLM benchmarks explainer covers why benchmark numbers transfer poorly, and the LLM leaderboards guide explains how to read boards like this one.

How to set up LLM hallucination detection

  1. Decide which failure matters. For RAG and summarization, focus on faithfulness to the provided context. For open-ended answers, focus on factual errors and uncertainty.
  2. Build a labeled set. Collect 50 to 200 real outputs and have domain experts mark each as supported or not. This is your ground truth.
  3. Pick a detector and calibrate it. Run a faithfulness metric or specialist model on the labeled set and measure agreement with your experts before using its scores.
  4. Set a threshold and route. Decide what score triggers a block, a warning, a fallback answer or human review.
  5. Monitor in production. Score a sample of live traffic and log the results next to traces, using one of the LLM observability tools.
  6. Fix the cause, not just the symptom. Many hallucinations come from bad retrieval, so check context recall and precision before tuning prompts.

Pros and cons of automated detection

Pros: catches many unsupported claims cheaply, works in CI and in production, and gives a trend line for quality over time.

Cons: judges and detectors make their own mistakes, sampling methods multiply cost and latency, and no method reliably detects confident, consistent errors without an external source.

Who should use which approach

Teams running RAG or summarization should start with a faithfulness metric in DeepEval, Ragas or TruLens, and consider HHEM for cheaper high-volume scoring. Teams with open-ended assistants should add sampling-based uncertainty on high-stakes questions and route uncertain answers to a safer response. Enterprises wanting a managed option can look at hosted evaluation platforms such as Patronus AI or Galileo, which Cisco acquired in 2026. Everyone should keep a human-labeled set to keep the detectors honest; our playbook on how to evaluate LLM apps shows how.

FAQ

How do you detect hallucinations in LLMs?

If you have source documents, split the answer into claims and check each against the sources with an LLM judge or a detector model such as HHEM. Without sources, sample several answers and measure disagreement with methods like SelfCheckGPT or semantic entropy, or look claims up in a trusted knowledge base.

What is the best tool for hallucination detection?

For RAG applications, the faithfulness metrics in DeepEval, Ragas and TruLens are the most common open-source choices, and Vectara's open HHEM model is a cheaper alternative at volume. Hosted platforms such as Patronus AI offer managed judge models.

Can hallucinations be detected without a reference?

Partly. Sampling-based methods such as SelfCheckGPT and semantic entropy flag answers where the model is uncertain, but they miss errors the model makes consistently. External fact checking against search or a knowledge base fills some of that gap.

What is the Vectara hallucination leaderboard?

It is a public leaderboard that measures how often language models add unsupported information when summarizing documents, scored with Vectara's open HHEM evaluation model. It reports hallucination rate, factual consistency rate and answer rate.

What is semantic entropy?

Semantic entropy is a method, published in Nature in 2024, that samples several answers to a question, groups them by meaning, and measures how spread out those meanings are. High semantic entropy suggests the model is confabulating.

Can hallucinations be eliminated completely?

No current technique eliminates them. Better retrieval, grounding instructions, detection with fallbacks and human review for high-stakes outputs can reduce their frequency and impact.