Prompt
Can you help me design an observability workflow for an LLM app that has retrieval, tools, and a human review loop?
Latest observation
Yes — here’s a practical observability workflow you can use for an LLM app that includes retrieval, tool use, and a human review loop.
1) Start with the workflow you want to observe
A good LLM app trace usually looks like:
- User request
- Preprocessing / routing
- Retrieval
- Prompt assembly
- LLM call
- Tool calls as needed
- Post-processing
- Human review if confidence is low / policy requires it
- Final response
- Feedback capture
You want observability at each step, not just for final outputs.
2) Define the core objects to log
Use a consistent set of fields for every run, span, and event.
Run-level fields
run_idsession_iduser_idor anonymized identifiertenant_idtimestamp_start,timestamp_endapp_versionprompt_versionmodel_namelatency_totalstatus(success,error,review_required,rejected)cost_estimatetoken_input,token_output
Step/span-level fields
For each stage:
span_idparent_span_idstep_name(retrieval,tool_call,review, etc.)start_time,end_timelatencystatusinput_summaryoutput_summaryerror_type,error_messageif any
LLM-specific fields
prompt_template_idsystem_prompt_versionretrieved_context_idscontext_token_counttemperature,top_pmodel_providerfinish_reasonstructured_output_validif applicable
Retrieval fields
queryretriever_nameindex_nametop_kdocuments_returneddocument_idsscoresreranker_usedretrieval_latencycontext_overlapor dedup stats
Tool-use fields
tool_nametool_inputtool_outputtool_latencytool_statusretry_countexternal_dependencytimeout_flag
Human review fields
review_required_reasonreviewer_idreview_decision(approve,edit,reject)review_notesdiff_from_model_outputreview_latencypolicy_tags
Feedback fields
thumbs_up/downuser_reported_issueresolution_statusground_truth_availablelabel_source(user,reviewer,offline_eval)
3) Trace the whole request as a single graph
Use distributed tracing semantics:
- One trace per request
- Each major stage is a span
- Nested spans for sub-operations:
- retrieval → query embedding, vector search, rerank
- tool call → API request, API response, retries
- review → queue wait, human decision
This lets you answer:
- Where time is spent
- Which steps fail most
- Whether retrieval quality affects answer quality
- Whether human review is slowing throughput
4) Instrument the retrieval layer
Retrieval issues are often the root cause of bad outputs.
Track:
- query text or hashed/redacted version
- embedding model/version
- documents retrieved
- ranking scores
- reranker changes
- citation coverage
- whether retrieved docs were actually used in final response
Useful metrics:
- retrieval latency
- hit rate
- MRR / nDCG if you have labels
- answer groundedness
- context utilization
- duplicate context rate
Recommended checks:
- empty retrieval results
- low-score retrievals
- stale document versions
- context truncation
- irrelevant or contradictory retrieved docs
5) Instrument tool calls separately
Tool use should be observable as its own subsystem.
Track:
- tool name and version
- input/output payloads
- retries and fallback behavior
- timeouts
- rate limits
- failure categories
- side effects performed
Recommended safeguards:
- log tool intents before execution
- require idempotency keys for writes
- record whether the tool result was actually incorporated into the answer
- classify tools as read-only vs write-capable
Useful alerts:
- spike in tool failures
- slow external dependency
- unusual tool invocation frequency
- tool output schema drift
6) Build a human review loop with visibility
Human review should be treated as an explicit workflow state, not an ad hoc override.
Track:
- why review was triggered
- queue time
- reviewer assignment
- SLA breach
- reviewer edits vs full rejection
- disagreement rate between model and human
- reasons for escalation
Metrics:
- review rate
- approval rate
- edit rate
- reject rate
- average time in queue
- human/model disagreement rate
This helps you spot:
- prompts that are too risky
- tasks the model handles poorly
- reviewer bottlenecks
- policy gaps
7) Add quality evaluation at multiple levels
Don’t rely only on production feedback.
Online metrics
- latency
- error rate
- token usage
- tool failure rate
- review rate
- user satisfaction
Offline evals
- task success rate
- factuality / groundedness
- retrieval precision
- tool correctness
- policy compliance
- human agreement
LLM-specific evals
- hallucination rate
- citation correctness
- instruction-following score
- refusal accuracy
- structured output validity
You can run these:
- on sampled production traces
- on golden datasets
- after prompt/model changes
- after retrieval index updates
8) Create a debugging workflow for bad outputs
When something goes wrong, your workflow should answer:
- Was the user request understood correctly?
- Did retrieval fetch the right context?
- Did the LLM use the context properly?
- Did any tool fail or return bad data?
- Was the response modified by human review?
- Did policy or formatting constraints break the output?
A good debug view shows:
- full trace
- prompt and prompt version
- retrieved docs
- tool inputs/outputs
- intermediate reasoning artifacts if allowed
- reviewer edits
- final output
- feedback
9) Use sampling and redaction carefully
Because prompts and tool inputs may contain sensitive data:
- redact PII
- hash or tokenize user identifiers
- store full payloads only in secure environments
- sample verbose traces
- separate audit logs from debug logs
Recommended pattern:
- always store metadata
- store content selectively
- retain only what you need for debugging and evaluation
10) Set alerts and dashboards
Dashboards
- request volume
- latency percentiles
- error rates
- retrieval quality
- tool success rates
- review throughput
- cost per request
- human override rates
Alerts
- p95 latency > threshold
- review queue growing too large
- retrieval empty-result spike
- tool failure spike
- model output schema failure
- cost anomaly
- sudden drop in user satisfaction
11) A simple reference architecture
You can structure it like this:
Application layer
- request router
- prompt builder
- response formatter
Observability layer
- trace manager
- span logger
- metrics emitter
- event bus
Subsystems
- retrieval service
- tool execution service
- human review queue
- evaluation service
Storage
- trace store
- metrics store
- audit log store
- redacted transcript store
Analysis
- dashboards
- anomaly detection
- offline eval jobs
- trace replay
12) Minimal event schema example
Here’s a compact JSON-style event model:
{
"run_id": "r123",
"session_id": "s456",
"span_id": "sp1",
"parent_span_id": null,
"step_name": "retrieval",
"status": "success",
"start_time": "2026-10-09T10:00:00Z",
"end_time": "2026-10-09T10:00:01Z",
"metadata": {
"query": "How do I reset my account?",
"retriever_name": "vector_search_v2",
"top_k": 5,
"document_ids": ["d1", "d9", "d14"],
"scores": [0.91, 0.84, 0.72]
}
}
You’d create similar spans for:
- prompt assembly
- LLM call
- each tool call
- human review
13) Recommended operating loop
A practical weekly loop:
- Review top failure traces
- Sample low-confidence requests
- Compare retrieval quality vs final answer quality
- Inspect tool errors and timeouts
- Review human edits/rejections
- Run offline evals on new prompts/models
- Ship improvements
- Track metric changes after release
14) If you want a strong baseline, prioritize these 5 things
- End-to-end tracing
- Retrieval visibility
- Tool call observability
- Human review metrics
- Offline evaluation tied to production traces
If you want, I can turn this into one of these:
- a concrete architecture diagram
- a JSON/SQL schema for logs and traces
- a dashboard/alert spec
- a LangChain / OpenTelemetry style implementation plan