Prompt
Can you help me figure out what observability setup I need for an LLM app that has retrieval, tool calls, and human review?
Latest observation
Yes. For an LLM app with retrieval, tool calls, and human review, you want observability at three levels:
- Request / conversation level — what happened for each user interaction
- Component level — what happened inside retrieval, tools, and review
- System level — trends, failure rates, latency, cost, quality, and safety
Here’s a practical setup.
1) What you need to observe
A. End-to-end traces
You want a trace for every user request that shows:
- user input
- prompt(s) sent to the model
- model output
- retrieved documents / chunks used
- tool calls made
- tool outputs
- human review decisions
- final response returned
This gives you a single timeline of the entire workflow.
B. Metrics
Track aggregate metrics such as:
- request volume
- latency by stage:
- retrieval latency
- model latency
- tool latency
- human review latency
- token usage
- cost per request
- retrieval hit rate
- tool success/error rate
- human override rate
- escalation rate
- rejection rate
- hallucination / grounding failure rate
- safety policy violation rate
C. Logs / artifacts
Store the actual artifacts needed for debugging and audits:
- prompts and completions
- retrieved passages
- tool inputs/outputs
- citations
- reviewer comments
- model version / prompt version / retrieval config version
- user/session IDs, with privacy controls
D. Quality signals
You also need signals that approximate quality:
- user feedback
- reviewer feedback
- task success/failure
- groundedness / citation correctness
- tool result correctness
- answer completeness
- policy compliance
2) Observability for each part of the pipeline
Retrieval
For retrieval, instrument:
- query text
- embedding model version
- index / corpus version
- top-k results
- similarity scores
- reranker scores if used
- chunk IDs and source metadata
- whether retrieved content was actually used in the final answer
Useful metrics:
- recall proxy
- MRR / nDCG if you have labels
- no-result rate
- irrelevant retrieval rate
- retrieval latency
Tool calls
For each tool call, capture:
- tool name
- arguments
- timestamps
- response payload
- status / error
- retry count
- timeout count
- downstream dependency used
- whether the model used the result correctly
Useful metrics:
- call success rate
- timeout rate
- error rate by tool
- avg latency by tool
- retries per request
- tool-induced failure rate
Human review
For human-in-the-loop review, capture:
- what was sent to reviewer
- reviewer decision
- reason code / rubric label
- time to review
- reviewer identity or role
- whether reviewer modified the output
- whether the model was accepted/rejected/escalated
Useful metrics:
- review queue depth
- median review time
- approval rate
- rejection rate
- inter-reviewer agreement
- escalation rate
- SLA breach rate
3) Recommended architecture
A good setup usually has:
Tracing
Use distributed tracing with spans for:
- user request
- retrieval
- LLM inference
- each tool call
- human review step
- final response
Each span should carry metadata like:
- trace ID
- request ID
- tenant ID
- model name/version
- prompt version
- retrieval index version
Metrics backend
Send numeric metrics to something like Prometheus, Datadog, Grafana, New Relic, etc.
Log store / event store
Keep structured event logs for:
- prompts/completions
- retrieval results
- tool events
- human review events
Evaluation pipeline
Run offline and online evals:
- golden set replay
- regression testing on prompts/models
- groundedness checks
- tool correctness checks
- red-team / safety evals
4) What to log in a single request trace
A minimal trace schema could include:
trace_idrequest_iduser_idor pseudonymous IDconversation_idtimestampinput_textretrieval: [ {doc_id, chunk_id, score, source} ]prompt_versionmodel_namemodel_versioncompletion_texttool_calls: [ {tool_name, args, result, latency_ms, error} ]human_review: {reviewed: true, decision, reviewer_role, latency_ms, comments}final_answerfeedbackcost_tokenssafety_flags
5) Alerts you should set up
Important alerts:
- model error rate spikes
- retrieval returns empty / low-confidence results
- tool error or timeout spikes
- review queue growing too long
- unusually high hallucination / rejection rate
- sudden cost increase
- latency p95/p99 regression
- prompt or model version changes causing quality drop
6) Privacy and compliance considerations
Because LLM apps often handle sensitive text:
- redact or tokenize PII before storing logs where possible
- separate raw content from metadata
- use role-based access controls
- limit retention of prompts and outputs
- store audit trails for human review
- encrypt data at rest and in transit
If regulated, consider:
- retention policies
- access logging
- data minimization
- reviewer accountability
- explainability/citation tracking
7) A practical “minimum viable” setup
If you want the smallest useful version, start with:
Must-have
- distributed tracing across retrieval, model, tools, and review
- structured logs of prompts, retrievals, tool calls, and review decisions
- latency and error metrics
- token/cost tracking
- human review SLA tracking
Nice-to-have next
- automated evals for groundedness and tool correctness
- feedback loop from reviewers and users
- experiment tracking for prompts/models
- quality dashboards by workflow stage
8) Suggested dashboard layout
A good dashboard would show:
Top-level
- total requests
- success rate
- average latency
- cost/request
- review rate
Workflow breakdown
- retrieval success rate
- tool success rate
- review approval/rejection rate
- escalation rate
Quality
- grounded answer rate
- citation coverage
- user satisfaction
- reviewer agreement
Reliability
- p95 latency by stage
- timeout/error counts
- queue backlog
- retries
9) If you want to implement this quickly
A common stack looks like:
- Tracing: OpenTelemetry
- Metrics: Prometheus + Grafana, or Datadog
- Logs: structured JSON logs to ELK/Datadog/Splunk
- LLM-specific evals: custom evaluation jobs or tools like LangSmith / Arize Phoenix / WhyLabs / Humanloop depending on your stack
- Human review workflow: internal app with audit logs
10) Best practice: tie everything to a single request ID
The most important design choice is to assign a single trace/request ID that follows the request through:
- retrieval
- model calls
- tool calls
- review
- final response
That’s what makes debugging possible.
If you want, I can turn this into one of these next:
- a concrete observability architecture diagram
- a list of exact metrics/events to instrument
- a sample trace schema in JSON
- a tooling recommendation by stack (Python, OpenAI API, LangChain/LlamaIndex, etc.)