Prompt

How do I measure hallucinations in a chatbot response?

Artificial Intelligence / AI Developer Tools1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

To measure hallucinations in a chatbot response, you want to check whether the answer is grounded in trusted sources and whether any claims are unsupported, incorrect, or fabricated.

1) Define what “hallucination” means for your use case

A hallucination is usually any of these:

  • Unsupported claim: the chatbot states something not found in the source material
  • Incorrect claim: the statement conflicts with the source or known facts
  • Fabricated detail: a made-up name, date, citation, quote, or event
  • Overconfident uncertainty: the chatbot presents uncertain info as certain

If your chatbot is using retrieval (RAG), the key question is:

“Is each answer statement supported by the retrieved context?”


2) Use claim-level evaluation

Break the response into individual claims, then check each one against a reference source.

Example response:

“The product launched in 2022, supports SSO, and is available in Europe.”

Claims:

  1. Product launched in 2022
  2. Supports SSO
  3. Available in Europe

Label each claim as:

  • Supported
  • Contradicted
  • Not enough evidence

Then compute:

  • Hallucination rate = contradicted or unsupported claims / total claims
  • Faithfulness rate = supported claims / total claims

3) Common metrics

For fact-grounded systems

  • Precision of factual claims: how many statements are correct
  • Faithfulness / groundedness: how much of the answer is supported by the provided context
  • Hallucination rate: fraction of unsupported or false claims
  • Citation accuracy: whether cited sources actually support the claim

For retrieval-based systems

  • Context precision: how much retrieved context is relevant
  • Context recall: whether the retrieved context contains the needed facts
  • Answer faithfulness: whether the answer uses only retrieved facts

4) How to measure it in practice

A. Human evaluation

Have annotators rate each response:

  • Supported
  • Unsupported
  • Contradicted
  • Unclear

This is the most reliable method, especially for nuanced answers.

B. Automated evaluation

Use:

  • LLM-as-judge: another model checks if each claim is supported by context
  • NLI/entailment models: determine whether context entails the answer
  • Fact-checking pipelines: extract claims and verify them against sources

Automated methods are cheaper, but less reliable than humans.

C. Reference-based testing

If you have a gold answer:

  • Compare the chatbot response to the reference
  • Measure semantic similarity, factual overlap, or exact correctness

This works well for structured tasks, but less so for open-ended chat.


5) A simple scoring scheme

You can score each response like this:

  • 0 = fully grounded, no unsupported claims
  • 1 = minor hallucination, mostly correct
  • 2 = moderate hallucination, several unsupported claims
  • 3 = severe hallucination, largely fabricated

Or use percentages:

  • 100% supported
  • 75% supported, 25% unsupported
  • etc.

6) Important caveats

  • A response can be useful but partially hallucinated
  • Some claims may be true in the real world but not supported by your provided source
  • Ambiguous questions make evaluation harder
  • “Hallucination” depends on the required standard: general knowledge vs. source-grounded answers

7) Practical recommendation

If you want a straightforward method:

  1. Split response into atomic claims
  2. Compare each claim to trusted source material
  3. Label claims as supported / unsupported / contradicted
  4. Compute hallucination rate
  5. Track results over many conversations

If you want, I can also give you:

  • a rubric for human annotators
  • a Python script to score hallucinations
  • or a prompt for an LLM judge to evaluate responses automatically.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.