Prompt
How do I monitor latency and errors on a model endpoint?
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
To monitor latency and errors on a model endpoint, the standard approach is to combine request-level logging, metrics, and alerts.
What to track
Latency
Track timing at several points:
- End-to-end request latency: time from client request received to response sent
- Model inference latency: time spent actually running the model
- Queue / preprocessing / postprocessing latency: if applicable
- Useful percentiles:
- p50: typical latency
- p95 / p99: tail latency, which often matters most in production
Errors
Track:
- HTTP error rates: 4xx, 5xx
- Inference failures: model exceptions, timeouts, bad input, OOMs
- Timeouts
- Invalid request rates
- Empty / malformed responses
- Retries and retry failures
How to implement monitoring
1. Add structured logging
Log each request with fields like:
- request ID / trace ID
- timestamp
- model name/version
- latency
- status code
- error type
- input/output sizes
- client/app identifier
Example:
{
"request_id": "abc123",
"model": "my-model-v1",
"latency_ms": 87,
"status": 200,
"error": null
}
2. Expose metrics
Publish metrics to a monitoring system such as:
- Prometheus + Grafana
- Datadog
- CloudWatch
- New Relic
- OpenTelemetry
Key metrics:
request_countrequest_latency_mshistogramerror_counttimeout_count5xx_rate4xx_rate
Use histograms or summary metrics for latency so you can compute p95/p99.
3. Add tracing
Use distributed tracing so you can see where time is spent:
- API gateway
- authentication
- preprocessing
- model inference
- postprocessing
- downstream calls
OpenTelemetry is a common choice.
4. Set alerts
Create alerts for:
- p95 latency above threshold
- error rate above threshold
- timeout spikes
- sudden drop in request volume
- model-specific failure spikes
Example alerts:
- p95 latency > 500 ms for 5 minutes
- 5xx error rate > 1% for 10 minutes
- timeout count increases by 3x baseline
5. Monitor by segment
Break down metrics by:
- model version
- endpoint
- region
- customer/app
- request type
- input size
This helps identify whether issues are isolated or widespread.
Best practices
- Use request IDs to correlate logs and traces
- Record baseline performance before setting alerts
- Monitor both overall availability and tail latency
- Separate client errors from server errors
- Include timeouts as errors
- Track resource metrics too:
- CPU
- memory
- GPU utilization
- queue depth
Simple checklist
If you want the minimum viable setup:
- Log every request with latency and status
- Export latency histogram and error counters
- Build a dashboard with p50/p95/p99 + error rate
- Set alerts for latency and 5xx spikes
- Add tracing if debugging becomes difficult
If you want, I can also show you:
- a Prometheus/Grafana setup
- Python/FastAPI example middleware
- or how to do this specifically for a cloud-hosted model endpoint.