Prompt
How do I measure whether AI answers are using my own content or skipping it?
Latest observation
To measure whether AI answers are using your own content vs skipping it, you want to evaluate grounding: how much of the answer is supported by your source material.
1) Define “using my content”
First decide what counts as using your content:
- Direct use: facts, phrasing, examples, or terminology from your docs appear in the answer
- Indirect use: the answer is paraphrased but still based on your content
- Skipping: the answer gives generic knowledge, misses key facts, or introduces unsupported claims
A practical definition:
An answer “uses your content” if most important claims can be traced to your sources.
2) Measure with simple metrics
You can score each answer against the relevant source document(s).
A. Coverage
How much of the answer is supported by your content?
- Break the answer into claims
- Mark each claim as:
- Supported
- Partially supported
- Unsupported
Then compute:
- Support rate = supported claims / total claims
B. Recall of key points
How many of your important source points did the AI include?
- Make a list of the key facts your content should surface
- Check whether the answer includes them
Example:
- Your content has 8 key facts
- The answer mentions 5
- Recall = 5/8
This tells you whether it’s skipping important content.
C. Precision
How many claims in the answer are actually grounded in your content?
- If the answer includes lots of generic or hallucinated info, precision is low
This tells you whether it’s making things up rather than using your content.
D. Quote/attribution rate
If your system supports citations:
- What fraction of answer sentences have citations to your content?
- How often are citations actually relevant?
3) Build a gold set
Create a small evaluation set of queries with expected source usage.
For each query, define:
- Relevant source docs
- Must-mention facts
- Nice-to-have facts
- Forbidden/unsupported facts
Then compare AI answers against this checklist.
Example template:
| Query | Must-use content | Expected answer behavior |
|---|---|---|
| “What is our refund policy?” | refund window, exceptions, process | should cite policy doc, not generic legal advice |
| “How do I reset my password?” | reset steps, escalation path | should follow internal help article |
4) Use retrieval-aware metrics if you have RAG
If your AI uses retrieval (RAG), measure both:
Retrieval quality
- Recall@k: did the retriever fetch the right doc chunks?
- MRR / nDCG: were the best chunks ranked highly?
Generation grounding
- Did the model actually use the retrieved text?
- Did it ignore the retrieved content and answer from memory?
This is important because a system can retrieve the right document but still skip it in the final answer.
5) Check for “content overlap” carefully
Don’t rely only on exact text overlap.
Exact matching can miss:
- paraphrases
- synonyms
- restructured explanations
Better methods:
- semantic similarity between answer and source chunks
- claim extraction + entailment checking
- citation validation
A good approach is:
- Split answer into atomic claims
- Compare each claim to source content
- Score whether the source entails the claim
6) Useful automatic evaluation methods
If you want something more scalable:
Heuristic scoring
- keyword overlap
- citation presence
- section coverage
LLM-as-judge
Ask another model:
- “Which parts of this answer are supported by the provided sources?”
- “Which key facts from the source are missing?”
- “Does the answer rely on outside knowledge?”
This works well if you use a strict rubric.
Entailment / NLI
Use an entailment model to test whether the source supports each claim.
7) A simple practical rubric
For each answer, score 0–2 on each dimension:
- Groundedness
- 0 = mostly unsupported
- 1 = mixed
- 2 = well grounded
- Coverage
- 0 = misses key source info
- 1 = partial
- 2 = complete
- Specificity
- 0 = generic
- 1 = somewhat specific
- 2 = clearly tied to your content
Then average scores across queries.
8) Red flags the AI is skipping your content
Look for:
- generic explanations that could apply anywhere
- answers that avoid mentioning your product/policy/terminology
- no citations when citations are expected
- answer quality stays similar even when your source docs change
- answer uses public knowledge instead of internal wording or constraints
9) Best overall method
If you want the most reliable measurement, use this pipeline:
- Create a benchmark of queries
- Annotate the source facts that should appear
- Generate answers
- Split answers into claims
- Score each claim for support
- Report
- support rate
- key-fact recall
- unsupported claim rate
- citation accuracy
That gives you a clear view of whether the AI is actually using your content.
If you want, I can also give you:
- a spreadsheet template for evaluating answers,
- a Python script to score grounding,
- or a rubric for RAG systems specifically.