Prompt
I'm building with OpenAI and need to see which requests are causing hallucinations
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
Yes — the practical way to debug hallucinations is to instrument each request so you can correlate bad outputs with the exact input, model settings, retrieved context, and any tool calls.
What to log for every request
At minimum, store:
- Request ID / trace ID
- Timestamp
- User/session ID (if applicable)
- Model name and version
- System prompt
- Developer prompt
- User prompt
- Tool/function definitions
- Retrieved context (RAG chunks, search results, citations)
- All tool calls and tool outputs
- Sampling settings:
temperature,top_p,max_output_tokens - Conversation history up to that turn
- Final model output
- Token usage
- Any safety refusals or errors
This lets you later answer:
- Did hallucinations happen mostly with low-quality retrieved context?
- Are they tied to a specific prompt template?
- Do they spike with a certain temperature or model?
- Do they happen after certain tool failures/timeouts?
Best practice: add a trace ID
When you send the request, include a unique ID and persist it in your app logs:
trace_id = str(uuid.uuid4())
Store that ID alongside:
- request payload
- response payload
- tool results
- user feedback / human review
What to compare when hallucinations happen
For each bad output, compare against:
- Prompt version
- Model version
- Retrieved documents
- Tool availability
- Conversation history length
- Temperature/top_p
- Latency and tool errors
If you use the OpenAI API
You can log the raw request/response objects from your app layer. If you’re using the OpenAI SDK, wrap the call and capture:
- inputs you send
- response text
- usage metadata
- any tool call structures
Stronger debugging setup
A good setup is:
- App logs for raw request/response
- Eval dataset of known bad cases
- Human review labels: hallucination / not hallucination / incomplete / unsupported
- Dashboards by model, prompt version, and retrieval quality
Helpful debugging questions
If you share more details, I can help you design the logging schema:
- Are you using RAG, tools, or just plain chat?
- Which SDK/language are you using?
- Do you want a database schema or a logging middleware example?
- Do you want to detect hallucinations automatically, or just trace which requests correlate with them?
If you want, I can give you a drop-in logging wrapper for Python or Node.js that records everything needed to trace hallucinations.