Prompt

I'm building a RAG assistant and want to track citations, hallucinations, and retrieval quality

Technology / Observability1 observationLast seen Oct 9, 2026

Latest observation

Oct 9, 2026OpenAI APIWeb search: off

A good way to think about this is to instrument your RAG system at three layers:

  1. Retrieval quality — did you fetch the right evidence?
  2. Grounding / hallucination — is the answer supported by that evidence?
  3. Citations — can you attribute each claim to a source chunk?

Below is a practical framework you can implement.


1) What to log for every RAG request

Store a single “trace” record per user query with:

  • query
  • retrieved_chunks: ids, text, source metadata, rank, scores
  • prompt_sent_to_llm
  • model_output
  • citations_emitted: claim → chunk ids
  • final_answer
  • latency breakdown
  • user feedback if available

A useful unit is a claim-level trace, not just a whole-answer trace.

Suggested trace schema

{
  "request_id": "uuid",
  "timestamp": "2026-10-09T12:34:56Z",
  "query": "What is the refund policy?",
  "retrieval": {
    "top_k": 5,
    "chunks": [
      {
        "chunk_id": "doc12#p3",
        "source": "help_center/refunds.md",
        "rank": 1,
        "score": 0.87,
        "text": "Refunds are available within 30 days..."
      }
    ]
  },
  "generation": {
    "answer": "You can request a refund within 30 days ...",
    "claims": [
      {
        "claim_id": "c1",
        "text": "Refunds are available within 30 days.",
        "citations": ["doc12#p3"]
      }
    ]
  },
  "metrics": {
    "retrieval_recall_at_k": 0.8,
    "faithfulness_score": 0.92,
    "citation_precision": 1.0
  }
}

2) Citation tracking: make the model cite at claim level

To avoid “one citation for the whole paragraph,” require the model to emit:

  • Answer text
  • Structured claims
  • Per-claim citations
  • Optional: quote spans from retrieved chunks

Example output format

{
  "answer": "Refunds are available within 30 days of purchase.",
  "claims": [
    {
      "text": "Refunds are available within 30 days of purchase.",
      "citations": [
        {
          "chunk_id": "doc12#p3",
          "quote": "Refunds are available within 30 days of purchase."
        }
      ]
    }
  ]
}

Best practices

  • Force citations on every factual sentence.
  • Prefer chunk ids over URLs alone.
  • Keep chunk text and offsets so you can verify exact support.
  • If unsupported, allow the model to say: “I couldn’t find evidence for that.”

3) Hallucination tracking: measure whether claims are supported

Hallucination in RAG usually means the answer contains claims that are not grounded in retrieved evidence.

Practical metrics

Track at the claim level:

  • Supported claim rate = supported claims / total claims
  • Unsupported claim rate = unsupported claims / total claims
  • Attribution accuracy = cited chunk actually supports the claim
  • Abstention rate = how often the model correctly says it doesn’t know

How to detect support

You can use a few approaches:

A. Rule-based / lexical overlap

Good for quick checks:

  • Extract named entities, numbers, dates, policy terms
  • Compare against retrieved chunk text

Pros: cheap, fast
Cons: weak for paraphrases

B. LLM-based entailment judge

Prompt a separate model:

  • “Does this evidence entail the claim?”
  • Output: supported / contradicted / not enough info

Pros: much better coverage
Cons: more cost, needs calibration

C. Embedding similarity

Useful as a signal, not a final decision:

  • claim embedding vs retrieved chunk embeddings

Pros: simple
Cons: not enough to prove support

Recommended hybrid

Use:

  • lexical checks for numbers/dates/entities
  • entailment judge for final support decision
  • similarity only as a retrieval diagnostic

4) Retrieval quality: measure whether the right evidence was found

Retrieval quality should be evaluated independently of generation.

Core metrics

At query level:

  • Recall@k: did any relevant chunk appear in top k?
  • MRR: how early did the first relevant chunk appear?
  • nDCG@k: ranking quality with graded relevance
  • Precision@k: how many retrieved chunks were actually relevant?
  • Coverage: are all needed subtopics represented?

If you have labeled relevant chunks

Great — compute classic IR metrics.

If you don’t have labels

Create an evaluation set with:

  • query
  • gold supporting chunk(s)
  • expected answer
  • maybe gold facts

Then use:

  • human labeling
  • LLM-assisted labeling with human review

RAG-specific retrieval diagnostics

Also track:

  • context utilization: did the model use the retrieved chunks?
  • evidence leakage: did it answer from prior knowledge instead of context?
  • chunk duplication: are top-k chunks redundant?
  • source diversity: are you retrieving from multiple docs when needed?

5) Suggested end-to-end evaluation pipeline

Offline eval

For a benchmark set:

  1. Retrieve top-k
  2. Generate answer
  3. Extract claims
  4. Judge each claim as supported / unsupported / contradicted
  5. Score citations and retrieval

Metrics to report

  • Retrieval: Recall@5, MRR, nDCG@10
  • Generation: faithfulness, correctness, completeness
  • Citation quality: citation precision, citation recall
  • Hallucination: unsupported-claim rate
  • Abstention: correct “I don’t know” rate

6) A simple scoring model you can implement

For each claim:

  • Supported = 1 if at least one retrieved chunk entails it
  • Cited correctly = 1 if the cited chunk supports it
  • Hallucinated = 1 if unsupported but asserted as fact

Then compute:

citation_precision = supported_and_cited / cited_claims
citation_recall = supported_and_cited / supported_claims
hallucination_rate = unsupported_claims / total_claims

You can also compute a per-answer score:

faithfulness = supported_claims / total_claims

7) Implementation tips

Prompting

Tell the model:

  • use only provided context
  • cite every factual statement
  • if evidence is insufficient, say so
  • avoid broad unsupported claims

Chunking

Better citations start with better chunks:

  • keep chunks semantically coherent
  • include headings and metadata
  • avoid overly large chunks
  • preserve document structure

Logging

Log raw retrieved text and the exact prompt, because debugging citation issues without them is painful.

Human review

Create a sampling dashboard for:

  • unsupported claims
  • low-confidence retrieval
  • answers with missing citations
  • disagreements between model and judge

8) Tooling ideas

You can implement this with:

  • OpenTelemetry for traces
  • LangSmith / Arize Phoenix / W&B / Weights & Biases for RAG observability
  • custom evaluation scripts for retrieval and faithfulness
  • LLM-as-judge for claim support

9) Recommended minimal setup

If you want the smallest useful version:

  1. Log query, retrieved chunks, answer, citations
  2. Extract claims from the answer
  3. Use an entailment judge to label each claim
  4. Compute:
    • retrieval Recall@k
    • supported-claim rate
    • citation precision
    • hallucination rate

That gives you a solid baseline.


If you want, I can also provide:

  • a Postgres schema for RAG traces,
  • a Python evaluation script,
  • or a LangChain/LlamaIndex instrumentation example.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.