Prompt
API endpoint monitoring with latency alerts
Latest observation
Oct 6, 2026
OpenAI APIWeb search: off
Here’s a practical pattern for API endpoint monitoring with latency alerts.
What to monitor
For each endpoint, track:
- Availability: success/failure rate, HTTP status codes
- Latency: p50 / p95 / p99 response times
- Throughput: requests per minute
- Error rate: 4xx / 5xx spikes
- Saturation: queue depth, CPU, memory, DB latency if applicable
Alerting strategy for latency
Use threshold + duration + percentiles to reduce noise.
Example alert rules:
- Warn if p95 latency > 300 ms for 5 minutes
- Critical if p95 latency > 800 ms for 2 minutes
- Page if p99 latency > 2 s for 1 minute on a core endpoint
Better yet, base alerts on:
- SLO burn rate: alert when latency SLO error budget is being consumed too quickly
- Per-endpoint thresholds: different APIs have different acceptable latencies
Example SLO
For a /checkout endpoint:
- SLO: 99% of requests under 500 ms
- Alert:
- Warning: >2x burn rate over 1 hour
- Critical: >4x burn rate over 5 minutes
Tooling options
You can implement this with:
- Prometheus + Alertmanager + Grafana
- Datadog
- New Relic
- CloudWatch + alarms
- OpenTelemetry + metrics backend
Prometheus example
Instrument endpoint latency as a histogram:
from prometheus_client import Histogram
REQUEST_LATENCY = Histogram(
"http_request_duration_seconds",
"HTTP request latency",
["method", "endpoint", "status"]
)
Then alert on p95 using PromQL:
histogram_quantile(0.95,
sum(rate(http_request_duration_seconds_bucket[5m])) by (le, endpoint)
) > 0.3
Example alert rule:
groups:
- name: api-latency
rules:
- alert: HighAPILatency
expr: histogram_quantile(0.95, sum(rate(http_request_duration_seconds_bucket[5m])) by (le, endpoint)) > 0.3
for: 5m
labels:
severity: warning
annotations:
summary: "High latency on {{ $labels.endpoint }}"
description: "p95 latency above 300ms for 5 minutes."
Recommended dashboards
Include:
- Latency by endpoint (p50/p95/p99)
- Error rate by endpoint
- Request volume
- Top slow requests
- Upstream dependency latency
Best practices
- Use histograms instead of simple averages
- Alert on symptoms, not just CPU/memory
- Separate alerts for user-facing critical endpoints
- Add synthetic probes from multiple regions
- Correlate alerts with deploys and dependency changes
If you want, I can also provide:
- a Prometheus/Grafana setup,
- a Datadog monitor configuration, or
- code examples for Node.js, Python, or Go.