Prompt
ChatGPT, how would you design webhook monitoring for delivery failures, latency, and duplicate events?
Latest observation
I’d design webhook monitoring as a combination of delivery telemetry, event lifecycle tracking, and alerting on SLOs/anomalies. The goal is to answer three questions reliably:
- Did the webhook get delivered?
- How long did it take end-to-end?
- Did we send or receive duplicates, and were they handled safely?
1) Track each webhook as a first-class delivery record
Create a webhook_delivery record for every attempted send with fields like:
delivery_idevent_iddestination_idattempt_numbercreated_atsent_atfirst_byte_atorresponse_received_atstatus(pending,success,retrying,failed,dead_lettered)http_statusresponse_body_hashor truncated responseerror_type(timeout,dns,tls,4xx,5xx,network)idempotency_keyrequest_size,response_time_ms
This gives you an auditable timeline and makes monitoring much easier.
2) Monitor delivery failures
What to measure
Track failure rates by:
- destination/customer
- event type
- status code class (
4xx,5xx, timeout) - retry attempt
- region/provider
- time window
Metrics
Useful counters and rates:
webhook_deliveries_totalwebhook_delivery_failures_totalwebhook_delivery_timeouts_totalwebhook_delivery_retries_totalwebhook_delivery_dead_lettered_total- failure rate = failures / attempts
- permanent failure rate = non-retryable failures / attempts
Alerts
Trigger alerts when:
- failure rate exceeds threshold over a rolling window
- a specific destination has repeated failures
- dead-letter queue starts growing
- retries spike sharply
- timeout rate increases
Practical rule
Separate failures into:
- transient: retryable (
5xx, timeout, network) - permanent: likely bad endpoint or auth issue (
401,403,404, malformed payload issues)
That distinction matters because transient failures shouldn’t page the team immediately unless sustained.
3) Monitor latency
You want both delivery latency and processing latency.
Key latency metrics
queue_delay_ms: time from event creation to dispatchrequest_duration_ms: time spent waiting for webhook responseend_to_end_latency_ms: event creation to successful delivery acknowledgmentretry_backoff_delay_ms: time spent waiting between retries
What to watch
- p50/p95/p99 latency
- tail latency per destination
- latency by event type
- latency under retry conditions
Alerts
- p95 latency exceeds SLO
- p99 latency spikes
- queue delay grows steadily
- delivery attempts are getting slower over time
Important distinction
A webhook can be “successful” but still too slow. If your SLO is “99% delivered within 60 seconds,” measure that directly rather than only looking at response time.
4) Detect duplicate events
Duplicates are usually unavoidable in webhook systems, so the design should assume at-least-once delivery.
Send-side protections
Include:
- globally unique
event_id idempotency_keyordelivery_id- monotonically increasing sequence if useful per resource
Receive-side handling
Consumers should:
- store processed
event_ids - reject/ignore duplicates
- make downstream actions idempotent
Duplicate monitoring
Measure:
- duplicate delivery rate
- duplicate event IDs received
- replay count per destination
- dedupe hits vs misses
Alerting
Alert if:
- duplicate rate exceeds normal baseline
- a destination starts reprocessing old events
- the same event is delivered many times due to retry failures
5) Use structured logs and distributed tracing
Logs
Log every attempt in structured format:
event_iddelivery_iddestinationattemptstatuslatency_mserrortrace_id
Tracing
Attach trace context so you can follow:
- event creation
- queueing
- dispatch
- destination response
- retries
This helps distinguish whether problems are:
- your queue
- your network
- the customer endpoint
- your retry logic
6) Dashboards to build
A good webhook dashboard should show:
Delivery health
- attempts/sec
- success rate
- failure rate
- retry rate
- dead-letter count
Latency
- p50/p95/p99 end-to-end latency
- p95 response time
- queue delay
Duplication
- duplicate events received
- dedupe success rate
- repeated delivery attempts by destination
Breakdown filters
- by customer
- by event type
- by region
- by endpoint host
- by status code
7) Retry and dead-letter strategy
Monitoring and retry strategy go together.
- exponential backoff with jitter
- cap retry count
- stop retrying on permanent failures
- move to dead-letter queue after max attempts
- expose dead-letter reasons
Monitor:
- retry success rate
- max-attempt exhaustion
- dead-letter growth
If dead letters aren’t reviewed, monitoring loses value, so set up a workflow to inspect and replay them safely.
8) Consumer-facing observability
If customers consume your webhooks, provide them with:
- delivery logs
- retry history
- event IDs
- timestamps
- failure reasons
- replay controls
This reduces support burden and helps customers debug their own webhook handlers.
9) SLOs and alerting philosophy
A simple SLO set might be:
- 99.9% of webhook deliveries succeed within 5 minutes
- 95% of deliveries complete within 10 seconds
- duplicate delivery rate below a small threshold, e.g. <0.1%
- dead-letter rate below threshold
Alert on:
- symptom-based issues that impact customers
- sustained degradation, not one-off blips
- destination-specific issues separately from global issues
10) Recommended implementation pattern
A solid architecture is:
- Event generated
- Persist event with unique
event_id - Enqueue delivery job
- Dispatch webhook attempt
- Record attempt metrics/logs
- Retry on transient failure
- Mark success/failure
- Deduplicate on both send and receive sides
If you want, I can also sketch:
- a schema for webhook delivery tables
- a Prometheus metric list
- or a Grafana dashboard layout for this monitoring setup.