Prompt
webhook retry backoff dead letter
Latest observation
Oct 5, 2026
OpenAI APIWeb search: off
A webhook retry / backoff / dead-letter design usually means:
- Retry failed webhook deliveries
- Use backoff to avoid hammering the receiver
- Move permanently failing events to a dead-letter queue (DLQ) for later inspection/replay
Typical flow
- Event occurs in your system
- You enqueue a webhook delivery job
- Worker sends the HTTP request
- If delivery succeeds: mark done
- If it fails:
- retry after a delay using backoff
- stop after max attempts or on non-retriable errors
- If still failing: send to DLQ
Retry rules
Common approach:
- Retry on transient failures:
- network errors
- timeouts
- 429 Too Many Requests
- 5xx responses
- Do not retry on permanent failures:
- 400 Bad Request
- 401/403 auth issues
- invalid endpoint config
- unsupported payload
Backoff strategies
1. Exponential backoff
Delay grows each attempt:
- attempt 1: 1s
- attempt 2: 2s
- attempt 3: 4s
- attempt 4: 8s
Usually add jitter to avoid thundering herd:
- randomize delay by ± some percentage
Example:
delay = base * 2^attempt + jitter
2. Linear backoff
Increase by a fixed amount:
- 5s, 10s, 15s, 20s
3. Capped backoff
Put a max delay:
- exponential growth up to 10 minutes max
Dead-letter queue
DLQ stores webhook attempts that failed too many times.
Store:
- event id
- webhook target URL
- payload
- last error
- attempt count
- timestamps
- next retry time
- correlation/request id
Use DLQ for:
- manual investigation
- replay after fixing receiver/config
- alerting on systemic failures
Important reliability concerns
- Idempotency: webhook receivers should dedupe using event id
- At-least-once delivery: most systems should assume duplicates
- Timeouts: keep request timeout short, e.g. 5–15 seconds
- Signing: sign payloads to verify authenticity
- Ordering: if ordering matters, retries can complicate things
- Concurrency control: avoid multiple workers delivering same webhook simultaneously
Example retry policy
- Max attempts: 8
- Base delay: 1s
- Backoff: exponential
- Max delay: 10 min
- Retryable status codes: 429, 500–599
- DLQ after 8 failures
Pseudocode
send_webhook(event):
for attempt in 1..max_attempts:
result = http_post(event.url, event.payload)
if result.success:
mark_delivered(event)
return
if result.status in [400,401,403,404]:
mark_dead_letter(event, reason="permanent_failure")
return
delay = min(base * 2^(attempt-1), max_delay)
delay = apply_jitter(delay)
sleep(delay)
mark_dead_letter(event, reason="max_attempts_exceeded")
Recommended production pattern
- Queue delivery jobs
- Use exponential backoff with jitter
- Classify retryable vs non-retryable errors
- Persist attempt history
- Send exhausted failures to DLQ
- Provide replay tooling
If you want, I can also give you:
- a database schema for webhook retry/DLQ,
- code examples in Python/Node/Go,
- or a system design diagram for this pipeline.