Prompt
How do I monitor latency, token spend, and timeout spikes in production for an LLM app?
Latest observation
To monitor an LLM app in production, track three layers: request performance, cost usage, and reliability/failures. The key is to measure them per model, per endpoint, per tenant/user, and per prompt type so spikes are actionable.
1) What to instrument on every LLM request
Log or emit metrics for each call:
- request_id / trace_id
- timestamp
- endpoint / feature
- model name + version
- provider
- prompt type / route
- tenant / customer / org
- input token count
- output token count
- total tokens
- latency:
- time to first token
- full completion time
- status:
- success / timeout / rate limit / context length error / provider error / cancelled
- retry count
- streaming vs non-streaming
- cache hit/miss if applicable
- estimated cost or token price tier
If you only do one thing: make sure every request has tokens + latency + status + model + tenant.
2) Latency monitoring
Track latency with percentiles, not just averages.
Metrics to chart
- P50, P95, P99 latency
- time to first token for streaming apps
- end-to-end request latency
- model inference latency if you can separate it from app overhead
- queue wait time if requests are buffered
Alerts to set
- P95 latency > baseline by, say, 30–50%
- P99 latency spikes above a hard threshold
- time to first token jumps suddenly
- latency increases for one model/tenant/region only
Good breakdowns
Slice latency by:
- model
- prompt/template
- tenant/customer
- region
- input token bucket size
- streaming/non-streaming
- provider
That helps determine whether spikes are due to larger prompts, provider degradation, or your own app logic.
3) Token spend monitoring
Token usage is your main cost driver.
Track
- input tokens
- output tokens
- total tokens
- tokens per request
- tokens per user / tenant / day
- cost per request
- cost per feature
- cost per successful outcome
Useful views
- daily spend trend
- top tenants by cost
- top endpoints by cost
- token distribution over time
- output-token inflation after a prompt change
Alerts
- daily spend exceeds budget
- tenant spend spikes unexpectedly
- average tokens/request rises sharply
- output token count grows after a release
- prompt retries increase spend
Important tip
If you can, record prompt version. Many cost regressions come from a prompt change that increases output verbosity or input size.
4) Timeout and failure spike monitoring
Timeouts often show up before broader outages.
Track failure categories separately
Don’t lump everything into “error.” Use:
- timeout
- provider error
- rate limit
- context length exceeded
- tool/function error
- parse/validation error
- cancelled by user
- retry exhausted
Alerts
- timeout rate above threshold, e.g. > 1–2%
- timeout rate doubles from baseline
- spike in rate limits
- spike in provider 5xxs
- retries increase sharply
- success rate drops for one model or region
Also monitor
- circuit breaker activations
- fallback rate to alternate model
- queue backlog
- concurrency saturation
- request abandonment
5) Dashboards to build
A practical dashboard set:
Executive / overview
- request volume
- success rate
- P95 latency
- timeout rate
- token spend today vs budget
- cost per successful request
Operations
- errors by type
- retries
- provider status
- queue depth
- active workers / concurrency
- regional breakdown
Cost
- input/output tokens
- cost by tenant
- cost by endpoint
- cost by model
- top expensive prompts
Performance debugging
- latency by model
- latency by input size
- latency by tenant
- first-token latency
- tail latency over time
6) Set up traces, not just metrics
Use distributed tracing so you can see the whole path:
- user request enters app
- prompt assembly
- retrieval step
- tool calls
- LLM call
- post-processing
- response return
This lets you answer:
- Is the delay in retrieval, tool execution, or the model call?
- Are timeouts caused by upstream API latency or your own code?
- Are retries hiding a provider issue?
7) Sample alert thresholds
Good starting points:
- P95 latency: alert if 30% above 7-day baseline for 10–15 minutes
- Timeout rate: alert if > 2x baseline or above 1%
- Token spend: alert if daily spend > 80% of budget
- Cost/request: alert if rises 20–30% week-over-week
- Retries: alert if retry rate doubles
- Rate limits: alert on sustained non-zero rate limits
Adjust thresholds to your app’s normal variance.
8) Practical stack options
You can implement this with:
- OpenTelemetry for traces/metrics
- Prometheus + Grafana for metrics dashboards
- Datadog / New Relic / Honeycomb for observability
- Cloud provider logs/metrics for infrastructure
- your data warehouse for cost analytics
For LLM-specific telemetry, many teams also use an application-level event schema and export to a warehouse.
9) Example metric schema
A simple event record per request:
{
"request_id": "abc123",
"tenant_id": "tenant_42",
"route": "summarize_document",
"model": "gpt-4.1",
"provider": "openai",
"prompt_version": "v12",
"input_tokens": 1820,
"output_tokens": 240,
"latency_ms": 1840,
"time_to_first_token_ms": 620,
"status": "success",
"error_type": null,
"retry_count": 1,
"estimated_cost_usd": 0.0214
}
10) Biggest mistakes to avoid
- only tracking averages instead of percentiles
- not separating timeout vs rate limit vs provider errors
- not logging token counts per request
- not tagging metrics by model/prompt/tenant
- measuring only the LLM call, not the full request path
- failing to track prompt version changes
- alerting on raw volume without baselines
If you want, I can give you:
- a production-ready metric schema,
- a Grafana/Datadog dashboard layout, or
- a code example in Python/Node for OpenTelemetry instrumentation.