Prompt
model API hosting latency monitoring
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
If you’re looking to monitor latency for an AI/model API you’re hosting, the usual approach is to track a few key metrics and send them to your observability stack.
What to monitor
- Request latency
- p50 / p95 / p99
- end-to-end time and model-only time if you can separate them
- Throughput
- requests/sec, tokens/sec
- Error rate
- HTTP 4xx/5xx, model failures, timeouts
- Queueing / saturation
- queue wait time, GPU/CPU utilization, memory usage
- Token metrics (if applicable)
- input/output tokens, total tokens per request
- Cold start latency
- especially for autoscaled or serverless deployments
Basic setup
- Instrument your API
- add timing around:
- request received
- preprocessing
- model inference
- postprocessing
- response sent
- add timing around:
- Export metrics
- Prometheus, OpenTelemetry, Datadog, CloudWatch, etc.
- Create dashboards
- latency percentiles over time
- latency by route/model/version/tenant
- Alert
- e.g. p95 > threshold for 5–10 min
- error rate spike
- queue time increasing
- Correlate with infrastructure
- GPU memory, CPU load, disk/network, pod restarts, autoscaling events
Example metrics names
request_duration_secondsinference_duration_secondsqueue_wait_secondsmodel_errors_totalrequests_total
Example Prometheus histogram
from prometheus_client import Histogram, Counter
REQUEST_LATENCY = Histogram(
"request_duration_seconds",
"Total request latency",
buckets=(0.05, 0.1, 0.2, 0.5, 1, 2, 5, 10)
)
ERRORS = Counter("model_errors_total", "Total model errors")
Useful alerts
- p95 latency > 2s for 10 minutes
- error rate > 1%
- queue wait > 500ms
- GPU utilization > 90% with rising latency
- cold starts increasing after deploy
If you want synthetic monitoring
Use an external checker to hit the endpoint every minute and measure:
- DNS/connect time
- TLS handshake
- time to first byte
- total response time
If you want, I can give you:
- a Prometheus/OpenTelemetry implementation example,
- a Grafana dashboard layout, or
- a ready-to-use latency alerting strategy for your model API.