Prompt
What should I use to debug LLM tool calls and retrieval failures?
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
To debug LLM tool calls and retrieval failures, use a combination of tracing, structured logs, and request/response inspection.
Best tools / approaches
-
Tracing / observability
- Use a tracing platform such as:
- LangSmith
- OpenTelemetry
- Arize / Phoenix
- Helicone
- These help you see:
- the exact prompt sent
- tool-call decisions
- tool inputs/outputs
- retrieval queries
- chunk selection and ranking
- latency and errors
- Use a tracing platform such as:
-
Structured logging
- Log:
- model name
- prompt/messages
- tool schema
- tool call arguments
- tool responses
- retrieved document IDs/chunks
- embedding/query text
- reranker scores
- Prefer JSON logs so they’re searchable.
- Log:
-
Replayable test cases
- Save failing conversations and retrieval queries.
- Build a small regression suite so you can rerun failures after fixes.
-
Prompt + tool schema inspection
- Many tool-call failures come from:
- ambiguous tool descriptions
- overly strict schemas
- bad examples
- missing required fields
- Inspect the exact tool definitions the model saw.
- Many tool-call failures come from:
-
Retrieval-specific debugging
- Check:
- chunking strategy
- embedding model quality
- top-k settings
- filters / metadata constraints
- query rewriting
- reranker behavior
- Compare:
- user query
- rewritten query
- retrieved docs
- final answer grounding
- Check:
Practical recommendation
If you want one stack to start with:
- LangSmith for LLM traces
- OpenTelemetry for app-wide tracing
- Structured JSON logs for backend debugging
What to look for first
- For tool calls: did the model choose the wrong tool, send malformed args, or ignore a tool entirely?
- For retrieval: did the right docs fail to retrieve, or were they retrieved but not used?
If you want, I can also give you a debugging checklist or a minimal tracing setup example for your stack.