Prompt
How do I monitor latency, token usage, and failures in production LLM workflows?
Latest observation
To monitor LLM workflows in production, track three layers: request performance, token/cost usage, and reliability/quality failures.
1) Latency: measure where time is spent
Track latency at multiple stages, not just end-to-end.
Key metrics
- End-to-end latency: total time from user request to final response
- Model latency:
- time to first token
- generation time
- total model round-trip time
- Pipeline latency:
- retrieval time
- tool/function call time
- reranker latency
- post-processing latency
- Queue/wait time: if requests are rate-limited or queued
Best practice
Break each workflow into spans:
- input preprocessing
- retrieval
- prompt assembly
- LLM call
- tool calls
- output validation
- response delivery
Then use tracing to identify bottlenecks.
Useful percentiles
Track:
- p50 for typical experience
- p95/p99 for tail latency
- max for severe outliers
2) Token usage: monitor cost and context efficiency
Tokens drive both cost and often latency.
Key metrics
- Input tokens
- Output tokens
- Total tokens
- Tokens by component:
- system prompt
- user prompt
- retrieved context
- tool outputs
- conversation history
- Tokens per request
- Tokens per successful task
- Context window utilization: how full the prompt is
- Prompt growth over time: detect drift in long conversations
Cost metrics
- cost per request
- cost per user/session
- cost per successful completion
- cost by feature, model, or tenant
Best practice
Log token counts for every call and aggregate by:
- model version
- workflow type
- customer/tenant
- prompt template version
This helps catch:
- prompt bloat
- retrieval overstuffing
- expensive model regressions
3) Failures: distinguish operational vs. task failures
Not all failures are API errors. In LLM systems, many failures are semantic.
Operational failures
- API timeouts
- rate limits
- network errors
- invalid responses / schema violations
- tool-call failures
- retries exhausted
- auth or quota issues
Task/quality failures
- hallucinations
- incorrect tool choice
- missing citations
- bad SQL / unsafe actions
- empty or irrelevant retrieval
- policy violations
- user dissatisfaction
Key metrics
- error rate
- timeout rate
- retry rate
- fallback rate
- structured output parse failure rate
- tool invocation success rate
- guardrail intervention rate
- human escalation rate
4) Add quality monitoring, not just uptime monitoring
For production LLMs, success means “did the workflow do the right thing?”
Quality signals
- user thumbs up/down
- task completion rate
- groundedness / citation coverage
- answer relevance
- factuality checks
- safety/policy flags
- agent success rate
- conversion or business KPI tied to the workflow
Recommended pattern
Use a mix of:
- online signals: user feedback, completion, retries
- offline evals: sampled prompts and outputs
- synthetic tests: fixed benchmark cases
- canary releases: compare old vs. new prompt/model
5) Instrumentation architecture
A good production setup usually includes:
Logs
Store:
- request id / trace id
- model name and version
- prompt template version
- token counts
- latency breakdown
- errors and exceptions
- tool calls
- final outcome
Metrics
Export to dashboards:
- latency percentiles
- token volume
- error rates
- fallback rates
- cost trends
Traces
Use distributed tracing across:
- app server
- retrieval service
- vector DB
- LLM provider
- external tools
Alerts
Alert on:
- latency spikes
- token spikes
- error spikes
- cost anomalies
- quality metric drops
6) Practical dashboard recommendations
A useful dashboard usually has:
Top row
- requests/min
- success rate
- p95 latency
- total token spend
- error rate
Breakdown charts
- latency by workflow stage
- tokens by model and prompt version
- failures by exception type
- retries/fallbacks over time
- cost by tenant or feature
Quality section
- task completion rate
- human rating
- output validation failures
- hallucination/grounding checks
7) Recommended alert thresholds
Start with thresholds based on baseline, not arbitrary numbers.
Examples:
- p95 latency increases by >30% over baseline
- token usage per request increases by >20%
- error rate exceeds 1–2%
- structured output parse failures exceed 0.5–1%
- fallback rate doubles
- cost per successful task rises materially
- user satisfaction drops below target
8) Common tools and patterns
You can implement this with:
- OpenTelemetry for traces/metrics
- Prometheus/Grafana for metrics dashboards
- Datadog/New Relic for APM
- LLM-specific observability tools like:
- LangSmith
- Arize Phoenix
- Helicone
- PromptLayer
- Weights & Biases Weave
9) Minimal checklist
If you want the shortest production-ready list, track:
- End-to-end latency
- p95/p99 latency
- Input/output tokens
- Cost per request
- Error rate
- Retry/fallback rate
- Tool-call success rate
- Output validation failure rate
- User feedback / task success
- Trace IDs across the workflow
If you want, I can also provide:
- a sample metrics schema
- a Grafana dashboard layout
- or a Python/TypeScript instrumentation example for LLM workflows.