Prompt
How do I evaluate whether an LLM answer is actually correct?
Latest observation
To evaluate whether an LLM answer is actually correct, use a mix of verification methods rather than trusting the surface quality of the response.
1) Check for the type of task
Different tasks need different validation:
- Factual questions: verify against reliable sources.
- Math/code: test the result directly.
- Reasoning/analysis: check whether the steps are valid and conclusions follow.
- Creative writing: “correctness” may mean consistency, not factuality.
2) Look for specific, checkable claims
Break the answer into atomic claims:
- dates
- definitions
- numbers
- named entities
- cause/effect statements
- quoted text
Then verify each claim independently. An answer can be mostly right but still contain one important falsehood.
3) Cross-check with authoritative sources
Use sources appropriate to the domain:
- General facts: reputable reference sites, textbooks, official documentation
- Medical/legal/financial: primary or professional sources, not just web summaries
- Programming: official docs, language specs, source code, tests
Prefer sources that are:
- current
- primary
- domain-specific
- consistent with one another
4) Test the answer in practice
For tasks where execution is possible:
- Code: run it
- Math: recompute independently
- Data analysis: reproduce with the same dataset
- Instructions: follow the steps and see if they work
If the answer predicts an outcome, check that the outcome actually occurs.
5) Inspect the reasoning, not just the conclusion
Even a correct-looking answer may be wrong if the logic is weak.
Ask:
- Are assumptions stated?
- Do the steps follow logically?
- Are there hidden leaps?
- Does it confuse correlation with causation?
- Does it overgeneralize from one example?
A solid answer should be traceable from premises to conclusion.
6) Watch for hallucination signals
Common warning signs:
- overly confident tone with no caveats
- specific-looking citations that may be invented
- unusual named entities, numbers, or dates
- answers that sound plausible but are hard to verify
- contradictions within the response
If it gives sources, verify that the sources actually exist and support the claim.
7) Ask the model to justify or self-check
Useful prompts:
- “Show the steps.”
- “What assumptions are you making?”
- “List the sources for each claim.”
- “What could make this answer wrong?”
- “Provide a concise verification checklist.”
This doesn’t prove correctness, but it can expose weak reasoning.
8) Use redundancy
Ask:
- the same question in a different way
- another model
- a human expert
- multiple sources
Agreement isn’t proof, but disagreement is a strong sign to investigate further.
9) Evaluate confidence separately from correctness
An LLM can sound confident while being wrong. Treat confidence as a clue, not evidence.
A better question is:
- “What evidence supports this?”
- “How easy would it be to verify?”
- “What is the failure mode if it’s wrong?”
10) For high-stakes domains, don’t rely on the LLM alone
For medical, legal, safety, or financial decisions:
- verify with qualified professionals
- consult primary sources
- use the model only as a drafting or brainstorming aid
A practical quick checklist
Before trusting an answer, ask:
- What exact claims is it making?
- Which claims can I verify?
- Do reliable sources support them?
- Does the logic actually work?
- Can I test it directly?
- Are there signs of hallucination?
- What happens if it’s wrong?
If you want, I can also give you:
- a short rubric for grading LLM answers,
- a workflow for fact-checking, or
- a checklist for coding/math answers specifically.