Prompt

ChatGPT, how would you design webhook monitoring for delivery failures, latency, and duplicate events?

Technology / API Platforms1 observationLast seen Oct 5, 2026

Latest observation

Oct 5, 2026OpenAI APIWeb search: off

I’d design webhook monitoring as a combination of delivery telemetry, event lifecycle tracking, and alerting on SLOs/anomalies. The goal is to answer three questions reliably:

  1. Did the webhook get delivered?
  2. How long did it take end-to-end?
  3. Did we send or receive duplicates, and were they handled safely?

1) Track each webhook as a first-class delivery record

Create a webhook_delivery record for every attempted send with fields like:

  • delivery_id
  • event_id
  • destination_id
  • attempt_number
  • created_at
  • sent_at
  • first_byte_at or response_received_at
  • status (pending, success, retrying, failed, dead_lettered)
  • http_status
  • response_body_hash or truncated response
  • error_type (timeout, dns, tls, 4xx, 5xx, network)
  • idempotency_key
  • request_size, response_time_ms

This gives you an auditable timeline and makes monitoring much easier.


2) Monitor delivery failures

What to measure

Track failure rates by:

  • destination/customer
  • event type
  • status code class (4xx, 5xx, timeout)
  • retry attempt
  • region/provider
  • time window

Metrics

Useful counters and rates:

  • webhook_deliveries_total
  • webhook_delivery_failures_total
  • webhook_delivery_timeouts_total
  • webhook_delivery_retries_total
  • webhook_delivery_dead_lettered_total
  • failure rate = failures / attempts
  • permanent failure rate = non-retryable failures / attempts

Alerts

Trigger alerts when:

  • failure rate exceeds threshold over a rolling window
  • a specific destination has repeated failures
  • dead-letter queue starts growing
  • retries spike sharply
  • timeout rate increases

Practical rule

Separate failures into:

  • transient: retryable (5xx, timeout, network)
  • permanent: likely bad endpoint or auth issue (401, 403, 404, malformed payload issues)

That distinction matters because transient failures shouldn’t page the team immediately unless sustained.


3) Monitor latency

You want both delivery latency and processing latency.

Key latency metrics

  • queue_delay_ms: time from event creation to dispatch
  • request_duration_ms: time spent waiting for webhook response
  • end_to_end_latency_ms: event creation to successful delivery acknowledgment
  • retry_backoff_delay_ms: time spent waiting between retries

What to watch

  • p50/p95/p99 latency
  • tail latency per destination
  • latency by event type
  • latency under retry conditions

Alerts

  • p95 latency exceeds SLO
  • p99 latency spikes
  • queue delay grows steadily
  • delivery attempts are getting slower over time

Important distinction

A webhook can be “successful” but still too slow. If your SLO is “99% delivered within 60 seconds,” measure that directly rather than only looking at response time.


4) Detect duplicate events

Duplicates are usually unavoidable in webhook systems, so the design should assume at-least-once delivery.

Send-side protections

Include:

  • globally unique event_id
  • idempotency_key or delivery_id
  • monotonically increasing sequence if useful per resource

Receive-side handling

Consumers should:

  • store processed event_ids
  • reject/ignore duplicates
  • make downstream actions idempotent

Duplicate monitoring

Measure:

  • duplicate delivery rate
  • duplicate event IDs received
  • replay count per destination
  • dedupe hits vs misses

Alerting

Alert if:

  • duplicate rate exceeds normal baseline
  • a destination starts reprocessing old events
  • the same event is delivered many times due to retry failures

5) Use structured logs and distributed tracing

Logs

Log every attempt in structured format:

  • event_id
  • delivery_id
  • destination
  • attempt
  • status
  • latency_ms
  • error
  • trace_id

Tracing

Attach trace context so you can follow:

  • event creation
  • queueing
  • dispatch
  • destination response
  • retries

This helps distinguish whether problems are:

  • your queue
  • your network
  • the customer endpoint
  • your retry logic

6) Dashboards to build

A good webhook dashboard should show:

Delivery health

  • attempts/sec
  • success rate
  • failure rate
  • retry rate
  • dead-letter count

Latency

  • p50/p95/p99 end-to-end latency
  • p95 response time
  • queue delay

Duplication

  • duplicate events received
  • dedupe success rate
  • repeated delivery attempts by destination

Breakdown filters

  • by customer
  • by event type
  • by region
  • by endpoint host
  • by status code

7) Retry and dead-letter strategy

Monitoring and retry strategy go together.

  • exponential backoff with jitter
  • cap retry count
  • stop retrying on permanent failures
  • move to dead-letter queue after max attempts
  • expose dead-letter reasons

Monitor:

  • retry success rate
  • max-attempt exhaustion
  • dead-letter growth

If dead letters aren’t reviewed, monitoring loses value, so set up a workflow to inspect and replay them safely.


8) Consumer-facing observability

If customers consume your webhooks, provide them with:

  • delivery logs
  • retry history
  • event IDs
  • timestamps
  • failure reasons
  • replay controls

This reduces support burden and helps customers debug their own webhook handlers.


9) SLOs and alerting philosophy

A simple SLO set might be:

  • 99.9% of webhook deliveries succeed within 5 minutes
  • 95% of deliveries complete within 10 seconds
  • duplicate delivery rate below a small threshold, e.g. <0.1%
  • dead-letter rate below threshold

Alert on:

  • symptom-based issues that impact customers
  • sustained degradation, not one-off blips
  • destination-specific issues separately from global issues

10) Recommended implementation pattern

A solid architecture is:

  1. Event generated
  2. Persist event with unique event_id
  3. Enqueue delivery job
  4. Dispatch webhook attempt
  5. Record attempt metrics/logs
  6. Retry on transient failure
  7. Mark success/failure
  8. Deduplicate on both send and receive sides

If you want, I can also sketch:

  • a schema for webhook delivery tables
  • a Prometheus metric list
  • or a Grafana dashboard layout for this monitoring setup.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as Circuit. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.