Prompt
How do I monitor latency and error rates across multiple model providers?
Latest observation
To monitor latency and error rates across multiple model providers, treat every provider call like a standard API dependency and instrument it the same way you would any critical service.
1) Capture the same core metrics for every request
Track these per provider, model, region, and endpoint:
- Request count
- Success count / error count
- Latency
- p50, p95, p99
- time to first token, if streaming
- total completion time
- Error rate
- HTTP 4xx / 5xx
- timeout rate
- rate-limit rate
- provider-specific API errors
- Retry count
- Token usage / cost
- Response size / streamed token throughput
At minimum, tag each metric with:
provider(OpenAI, Anthropic, Azure OpenAI, etc.)modeloperation(chat, embeddings, rerank, etc.)status(success/failure)error_typeregionordeploymentif relevant
2) Use distributed tracing around each model call
Wrap each provider call in a trace span. Record:
- start/end time
- request metadata
- provider/model tags
- retries
- error details
This lets you answer:
- Which provider is slow?
- Is latency caused by retries or the initial call?
- Are failures concentrated in one model or region?
Popular tools:
- OpenTelemetry
- Datadog APM
- New Relic
- Grafana Tempo + Prometheus
- Honeycomb
3) Centralize logs with structured fields
Log each request in a structured format, e.g.:
{
"provider": "anthropic",
"model": "claude-3.5-sonnet",
"latency_ms": 842,
"status": "error",
"error_type": "rate_limit",
"retry_count": 2,
"tokens_in": 1200,
"tokens_out": 340,
"request_id": "..."
}
This makes it easy to:
- compute provider-specific error rates
- correlate spikes with deployments
- investigate individual failures
4) Build dashboards by provider and model
Create dashboards showing:
- latency percentiles over time
- error rates by provider/model
- timeout and retry trends
- cost and token usage
- SLO compliance
Good views:
- Overall comparison table: provider, p95 latency, error rate, cost
- Time series: latency/error trends per provider
- Heatmap: model vs region vs error rate
- Top errors: by type and provider
5) Set SLOs and alerts
Define thresholds like:
- p95 latency > X ms for 5 minutes
- error rate > Y% for 10 minutes
- timeout rate > Z%
- rate-limit errors rising sharply
- one provider significantly worse than others
Alert on both:
- absolute thresholds
- relative degradation vs baseline or peers
Example:
- “Anthropic p95 latency is 2x the 24h baseline”
- “OpenAI 5xx rate exceeds 1%”
6) Add synthetic checks
Run scheduled test prompts against each provider:
- simple prompt
- medium prompt
- streaming prompt
- embeddings request
- tool/function-call request if applicable
Measure:
- availability
- latency
- consistency
- model regressions
This helps detect issues before users do.
7) Normalize provider-specific failures
Different providers report errors differently. Map them into a common taxonomy:
timeoutrate_limitauth_errorbad_requestserver_errornetwork_errorunknown
That makes cross-provider comparison much easier.
8) Sample implementation pattern
A good approach is to create a single wrapper for all provider calls:
- start timer / trace span
- call provider
- catch errors
- classify error
- emit metrics/logs
- propagate request ID / trace ID
This gives you one place to enforce observability across all vendors.
9) Recommended stack
A common setup:
- OpenTelemetry for traces
- Prometheus for metrics
- Grafana for dashboards/alerts
- Loki or ELK for logs
- Optional: Datadog or New Relic if you want an all-in-one platform
10) What to compare across providers
Useful comparative metrics:
- p50 / p95 / p99 latency
- success rate
- timeout rate
- rate-limit rate
- retry-adjusted success rate
- cost per successful request
- tokens per second
- user-visible failure rate
If you want, I can also give you:
- a concrete observability schema for model calls,
- a Python/TypeScript wrapper example, or
- a Grafana dashboard layout for multi-provider LLM monitoring.