Prompt
What should I use for webhook backoff and dead-letter handling?
Latest observation
Oct 8, 2026
OpenAI APIWeb search: off
For webhook delivery, a good default is:
Retry/backoff
Use exponential backoff with jitter.
Example policy:
- Immediate first retry only if the failure is clearly transient
- Then retry at:
1m, 5m, 15m, 1h, 6h, 24h - Add random jitter so many failed webhooks don’t retry at the same time
- Cap the max retry delay and total retry window
Why:
- Exponential backoff avoids hammering a failing receiver
- Jitter prevents retry storms
- It’s the most common and predictable pattern for webhook systems
Dead-letter handling
Move the webhook to a dead-letter queue / failed-delivery store after:
- max retry attempts reached, or
- retry window exceeded, or
- a non-retryable error is returned
Store enough data to inspect and replay:
- webhook payload
- target URL
- response status/body
- error type
- attempt count
- timestamps
- idempotency/event ID
Then provide:
- manual replay
- automatic replay after endpoint recovery
- alerting/monitoring for repeated failures
Practical recommendation
A solid default policy is:
- Retryable errors: network timeouts, 429, 5xx
- Non-retryable errors: 400/401/403/404/410, invalid payload, signature mismatch
- Max attempts: 8–12
- Backoff: exponential with full jitter
- DLQ: after max attempts or permanent failure
- Replay: admin-triggered and idempotent
If you want a simple implementation
Use:
- queue + retry metadata
- delivery worker
- dead-letter table/queue
- idempotency key per event
- per-destination rate limiting
If you tell me your stack (e.g. AWS/SQS, Kafka, Postgres, Redis, Celery, Temporal, etc.), I can suggest a concrete setup and retry schedule.