Prompt
How do I monitor webhook latency and failure rates?
Latest observation
Oct 5, 2026
OpenAI APIWeb search: off
To monitor webhook latency and failure rates well, track them at both the application level and the delivery pipeline level.
1) Measure the right metrics
Latency
Track:
- Request duration: time from webhook send to receiving a response
- End-to-end delivery latency: time from event creation to successful delivery
- Processing time: time your handler spends doing work
- Queue delay: if retries/queues are involved, time spent waiting before delivery
Useful percentiles:
- p50: typical latency
- p95 / p99: tail latency, where issues usually show up
Failure rates
Track:
- 4xx rate: usually bad payloads, auth, signature, or routing issues
- 5xx rate: server-side problems on your side
- Timeout rate
- Retry rate
- Permanent failure rate: events that ultimately never succeeded after retries
A common high-level metric:
- Success rate = successful deliveries / total attempts
- Also track success after retries separately from first-attempt success
2) Instrument both sender and receiver
On the sender side
Log or emit metrics for each attempt:
- event_id
- attempt number
- target URL
- status code
- response time
- timeout indicator
- retry reason
On the receiver side
Measure:
- request start/end time
- response code
- body validation failures
- signature verification failures
- business logic errors
3) Use structured logs
Include correlation fields so you can trace one event end-to-end:
event_iddelivery_idattemptproviderstatus_codelatency_mserror_typeretryable
Example log fields:
{
"event_id": "evt_123",
"attempt": 2,
"status_code": 500,
"latency_ms": 842,
"error_type": "db_timeout",
"retryable": true
}
4) Create dashboards
A good dashboard should show:
- Request volume
- Success rate
- 4xx/5xx breakdown
- Timeout count
- p50/p95/p99 latency
- Retry counts
- Top failing endpoints or event types
5) Set alerts
Alert on:
- Spike in 5xx or timeout rate
- Drop in success rate below threshold
- p95/p99 latency above SLO
- No deliveries / sudden traffic drop
- Repeated failures for a specific endpoint
Example alert thresholds:
- 5xx rate > 2% for 5 minutes
- p95 latency > 2s for 10 minutes
- timeout rate > 1% for 5 minutes
6) Handle retries carefully
Retries can hide real failure rates, so monitor:
- first-attempt failure rate
- final failure rate after retries
- retry exhaustion rate
This helps you distinguish:
- transient issues
- persistent misconfigurations
- endpoint outages
7) Add tracing if possible
Distributed tracing helps if webhook handling touches multiple services:
- API gateway
- auth/signature verification
- DB
- downstream APIs
Use trace IDs to follow one delivery through the system.
8) Tools commonly used
Depending on your stack:
- Metrics: Prometheus, Datadog, CloudWatch, Grafana
- Logs: ELK/OpenSearch, Loki, Splunk
- Tracing: OpenTelemetry, Jaeger, Tempo, Datadog APM
9) A practical minimal setup
If you want the simplest useful setup, start with:
- Counter for total deliveries
- Counter for successes, 4xx, 5xx, timeouts
- Histogram for latency
- Counter for retries and final failures
- Alert on error rate and p95 latency
If you want, I can also give you:
- a Prometheus/OpenTelemetry example, or
- a sample dashboard + alert thresholds for webhook monitoring.