Prompt
How do I monitor latency and errors for a model API?
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
To monitor latency and errors for a model API, track metrics at the request level and make them visible in logs, dashboards, and alerts.
1) Measure the right latency metrics
Collect these for every request:
- End-to-end latency: total time from request received to response returned
- Model inference latency: time spent inside the model call
- Queue/wait time: time waiting before execution
- Time to first token and tokens/sec if it’s a streaming LLM API
- P50 / P95 / P99 latency: percentiles are more useful than averages
2) Track errors by type
Record:
- HTTP status codes: 4xx, 5xx
- Timeouts
- Rate limit errors
- Validation/input errors
- Upstream/model errors
- Retries and retry exhaustion
- Cancelled requests
Also keep error rate = errors / total requests over time.
3) Add structured logging
Log each request with fields like:
- request_id / trace_id
- timestamp
- model name/version
- latency_ms
- status code
- error type/message
- input/output token counts
- user/org identifier if relevant
- retry count
Use JSON logs so they’re easy to search and aggregate.
4) Export metrics to a monitoring system
Common options:
- Prometheus + Grafana
- Datadog
- New Relic
- CloudWatch / Stackdriver / Azure Monitor
Recommended metrics:
request_latency_mshistogramrequest_counterror_counttimeout_countrate_limit_countretry_count
5) Use distributed tracing
If your API calls other services, use OpenTelemetry to trace:
- client → gateway → app → model provider → database/cache This helps identify where latency is introduced.
6) Set alerts
Alert on:
- P95 latency above threshold
- error rate above threshold
- timeout spikes
- sudden drop in throughput
- provider-specific failures
- retries increasing unusually
Example thresholds:
- P95 latency > 2s for 5 minutes
- error rate > 2% for 10 minutes
- timeout rate > 1% for 5 minutes
7) Monitor model-specific signals
For model APIs, also watch:
- token usage
- context length
- truncation rate
- streaming disconnects
- safety/filter rejections
- cache hit rate
- cold starts if using serverless/hosted inference
8) Create a simple dashboard
A good dashboard includes:
- Requests per minute
- Success/error rate
- Latency percentiles
- Top error types
- Latency by model/version
- Latency by region
- Token usage trends
9) Sample implementation pattern
In code, wrap the API call with timing and error capture:
import time
import logging
logger = logging.getLogger(__name__)
def call_model():
start = time.time()
try:
result = model_client.generate(...)
latency_ms = (time.time() - start) * 1000
logger.info("model_request", extra={
"status": "success",
"latency_ms": latency_ms,
"model": "my-model",
})
return result
except Exception as e:
latency_ms = (time.time() - start) * 1000
logger.exception("model_request_failed", extra={
"status": "error",
"latency_ms": latency_ms,
"model": "my-model",
"error_type": type(e).__name__,
})
raise
10) Practical checklist
- Add request timing
- Log all failures with error type
- Emit metrics histograms/counters
- Add tracing with request IDs
- Build dashboards for latency and error rate
- Set alerts
- Track model/version/token metrics
If you want, I can also give you:
- a Prometheus/OpenTelemetry example,
- a Python/FastAPI middleware, or
- a Datadog dashboard setup.