Prompt
How do I monitor and debug an LLM app in production?
Latest observation
Monitoring and debugging an LLM app in production is a mix of classic software observability and LLM-specific evaluation. The main idea is: log the right signals, define quality metrics, trace every request end-to-end, and build a feedback loop for failures.
1) What to monitor
Core system metrics
These are the usual production metrics:
- Request rate: traffic volume by endpoint / feature
- Latency: p50 / p95 / p99 end-to-end and per model call
- Error rate: API failures, timeouts, rate limits, malformed outputs
- Throughput: tokens/sec, requests/sec
- Cost: cost per request, per user, per workflow
- Retries / fallback usage: how often you retry or switch models
LLM-specific metrics
These matter more than traditional app metrics:
- Prompt size and completion size
- Token usage by request, user, tenant, feature
- Context window saturation: how close prompts get to max context
- Tool-call success rate: function calling / agent action success
- Structured output validity: JSON parse success, schema compliance
- Hallucination indicators: unsupported claims, missing citations, low retrieval grounding
- Retrieval quality if using RAG:
- retrieval hit rate
- relevance of retrieved chunks
- answer groundedness
- source coverage
- Conversation quality:
- user re-asks
- abandonment rate
- escalation to human
- thumbs up/down or other feedback
2) Instrument every LLM request
For each request, capture a trace like:
- request ID / trace ID
- user ID / tenant ID (if allowed)
- app feature / route
- model name + version
- system prompt version
- user prompt template version
- retrieved docs IDs
- tool calls and their results
- token counts in/out
- latency breakdown:
- prompt building
- retrieval
- model inference
- tool execution
- post-processing
- final output
- success/failure status
- safety/guardrail outcomes
- user feedback, if available
This lets you answer: What happened? Why? With which inputs? On which model/prompt version?
3) Use distributed tracing
Treat each LLM call as a span in a trace:
- parent span: user request
- child spans:
- retrieval
- reranking
- prompt assembly
- model call
- tool execution
- validation
- response formatting
With tracing you can find:
- which step is slow
- whether the model or your orchestration is the bottleneck
- whether failures cluster around a particular prompt version or tool
4) Log prompts and outputs carefully
You usually want to log:
- the exact prompt sent
- the exact model response
- intermediate tool outputs
- validation errors
But do this with privacy and security in mind:
- redact PII and secrets
- avoid storing sensitive user data unless necessary
- hash or tokenize identifiers
- use sampling for high-volume traffic
- set retention policies
If you can’t store full text, store:
- prompt template ID
- prompt hash
- key metadata
- selected excerpts
- embedding or fingerprint for clustering
5) Build quality checks in production
Add automated checks after generation:
- schema validation for JSON outputs
- policy checks for unsafe content
- citation checks for RAG answers
- grounding checks: does the answer reference retrieved evidence?
- business-rule validation: e.g. dates, totals, state transitions
- confidence thresholds: route low-confidence answers to fallback or human review
6) Detect drift and regressions
LLM apps can degrade because of:
- model version changes
- prompt edits
- retrieval corpus changes
- tool/API changes
- user behavior shifts
Track over time:
- latency
- cost
- answer acceptance
- task success
- retrieval relevance
- hallucination rate
- safety violations
When metrics move, compare:
- before/after prompt version
- before/after model version
- feature flags
- traffic segment
- tenant / locale / language
7) Create a debugging workflow
When something goes wrong, inspect in this order:
- Was the input bad?
- malformed prompt
- missing context
- wrong retrieved docs
- Did retrieval fail?
- irrelevant or stale chunks
- Did the model fail?
- poor reasoning
- instruction conflict
- truncation
- Did tool execution fail?
- API errors
- schema mismatch
- Did post-processing fail?
- parser bugs
- filtering too aggressive
A good trace should let you replay the exact path.
8) Use offline evals and golden sets
Production monitoring should be paired with offline evaluation:
- maintain a golden dataset of real tasks with expected outcomes
- run evals before every prompt/model change
- measure:
- task accuracy
- groundedness
- format validity
- refusal quality
- tool-use correctness
This is how you catch regressions before users do.
9) Add user feedback loops
Useful feedback signals:
- thumbs up/down
- “was this helpful?”
- correction submissions
- edits made by users
- time-to-resolution
- abandonment
These signals are noisy, but great for:
- ranking problematic prompts
- finding weak tasks
- prioritizing evals
10) Practical stack
A common stack looks like:
- Tracing: OpenTelemetry + your APM
- Logs: structured JSON logs
- Metrics: Prometheus / Datadog / Grafana
- LLM observability: LangSmith, Helicone, Arize Phoenix, Weights & Biases, Honeycomb integrations, etc.
- Eval pipelines: custom CI jobs + golden datasets
- Alerting: Slack/PagerDuty on spikes in error rate, cost, latency, or quality failures
11) Recommended alerts
Set alerts on:
- p95 latency jump
- completion error spike
- token cost spike
- JSON/schema failure rate
- tool-call failure rate
- retrieval empty-result rate
- user downvotes spike
- safety filter spike
- model fallback rate increase
12) Best practices
- Version prompts, tools, retrieval corpora, and models
- Make every request traceable
- Sample and redact logs
- Separate system bugs from model quality issues
- Test with real production-like prompts
- Track quality by segment, not just globally
- Have a fallback strategy:
- retry
- smaller model
- cached answer
- human handoff
A simple rule of thumb
If you can’t answer these 5 questions from your telemetry, improve observability:
- What exact input was sent?
- What exact model/version handled it?
- What retrieval/tool context was used?
- Where did latency or failure occur?
- Did the output actually satisfy the task?
If you want, I can also give you:
- a production observability checklist
- a recommended JSON log schema
- or a reference architecture for monitoring an LLM app.