Prompt
I'm building a production LLM workflow and need to monitor cost, latency, and errors
Latest observation
To monitor cost, latency, and errors in a production LLM workflow, set up observability at three levels: request, model, and business outcome.
1) What to measure
Cost
Track per request and aggregate:
- Input tokens
- Output tokens
- Total tokens
- Model name / version
- Cost per request
- Cost per user / tenant / feature
- Cost trends over time
Useful derived metrics:
- Cost per successful request
- Cost per 1k requests
- Cost per conversation/session
- Cost by prompt template or workflow step
Latency
Track:
- End-to-end request latency
- Time to first token if streaming
- Model inference latency
- Tool / retrieval / function-call latency
- Queue / retry delay
- P50 / P95 / P99 latency
Break latency down by:
- Prompt size
- Model choice
- Region
- Tool usage
- Retry count
Errors
Track:
- API failures
- Timeouts
- Rate limits
- Invalid inputs / schema mismatches
- Tool execution failures
- Hallucination-related downstream failures if you can detect them
- Retry exhaustion
- Partial stream failures
Also log:
- Error type
- Error message
- HTTP status / provider status
- Retry attempt
- Correlation ID / trace ID
2) Instrumentation design
Add a unique request ID
Generate a correlation ID for every workflow execution and pass it through:
- frontend
- backend
- LLM call
- retrieval/tool calls
- downstream services
This lets you trace a single user action end-to-end.
Log structured events
Use JSON logs with fields like:
request_iduser_idortenant_idworkflow_namemodelprompt_tokenscompletion_tokenslatency_msstatuserror_typecost_usdprompt_template_version
Separate spans for each step
If your workflow includes:
- retrieval
- reranking
- LLM generation
- tool execution
- post-processing
treat each as its own measured span so you know where time and failures occur.
3) Metrics dashboard
Build a dashboard with:
- Total cost today / this week / this month
- Cost by model
- Cost by feature
- Latency percentiles
- Error rate by provider/model
- Token usage distribution
- Retries per request
- Timeouts and throttling events
Good charts:
- Time series for cost and latency
- Stacked bar for cost by workflow step
- Heatmap for latency vs token count
- Pie/bar for error categories
4) Alerting
Set alerts on:
- Cost spikes over baseline
- P95 latency regressions
- Error rate > threshold
- Rate limit spikes
- Timeout spikes
- Unexpected token growth
- Model fallback frequency increasing
Example alert thresholds:
- Cost up > 20% day-over-day
- P95 latency > 2x trailing 7-day average
- Error rate > 3%
- Timeout rate > 1%
- Retry rate > 10%
5) Tracing and observability stack
A practical stack:
- OpenTelemetry for traces and metrics
- Prometheus + Grafana for metrics dashboards
- ELK / OpenSearch / Datadog / Honeycomb for logs and traces
- A cost calculator in your app to estimate spend from token usage and provider pricing
If you use OpenAI, Anthropic, or another provider, capture:
- model
- usage metadata
- finish reason
- request duration
- error payloads
6) Production best practices
Retry carefully
- Retry only transient failures
- Use exponential backoff
- Cap max retries
- Record retry count in metrics
Control spend
- Set token limits
- Choose smaller models for simple tasks
- Cache frequent prompts/results
- Use truncation and summarization for long contexts
- Route to cheaper models when possible
Sample verbose logs
To reduce observability cost:
- log everything for failures
- sample successful high-volume traffic
- keep full traces for a subset of requests
Redact sensitive data
Never log raw PII or secrets.
- redact user content where necessary
- store prompt hashes or partial content for debugging
- encrypt logs if they may contain sensitive text
7) Minimal schema to store per request
{
"timestamp": "2026-10-09T12:34:56Z",
"request_id": "req_123",
"user_id": "user_456",
"workflow": "support_reply",
"model": "gpt-4.1-mini",
"prompt_tokens": 1200,
"completion_tokens": 320,
"total_tokens": 1520,
"cost_usd": 0.0048,
"latency_ms": 1832,
"status": "success",
"error_type": null,
"retry_count": 0
}
8) If you want a simple implementation plan
- Add
request_ideverywhere - Log every LLM call with token counts, latency, status, and model
- Emit metrics to Prometheus/OpenTelemetry
- Build Grafana/Datadog dashboards
- Add alerts for cost, latency, and errors
- Review top-cost workflows weekly
- Add sampling and redaction
If you want, I can also give you:
- a reference architecture
- a Python/TypeScript instrumentation example
- a Grafana dashboard design
- or a production-ready logging schema for LLM calls