Prompt
I'm building a backend that needs retry and failover for email
Latest observation
Oct 5, 2026
OpenAI APIWeb search: off
For a backend that needs retry and failover for email, the most reliable pattern is:
Recommended approach
Use an email provider abstraction with:
- Primary provider (e.g. SES, SendGrid, Mailgun)
- Secondary provider for failover
- Queue + worker for retries
- Idempotency/deduplication to avoid duplicate sends
Basic flow
- Your app writes an email job to a queue.
- A worker sends the email through the primary provider.
- If it fails with a retryable error:
- retry with exponential backoff
- after N attempts, switch to the secondary provider
- Track delivery state in your DB.
Retry strategy
Retry only on transient failures, such as:
- timeouts
- rate limits (
429) - temporary 5xx provider errors
- network failures
Do not retry on permanent errors:
- invalid recipient
- blocked/blacklisted address
- malformed message
- authentication failures due to bad config
Suggested retry policy
- 1st retry: after 30s
- 2nd retry: after 2m
- 3rd retry: after 10m
- then fail over to secondary provider
- max total attempts: 5–7
Use exponential backoff with jitter.
Failover strategy
Two common options:
1) Sequential failover
Try provider A first, then provider B only if A fails.
- simpler
- best when you want one provider to handle most traffic
2) Weighted routing
Send 95% to A and 5% to B, with failover on error.
- better for resilience testing
- avoids cold-standby problems
Important: avoid duplicate emails
Retries and failover can cause duplicates if the first provider actually accepted the message but your app didn’t get the response.
To reduce this:
- generate a message id for each email job
- store send status in your DB
- make send operations idempotent in your own system
- if provider supports idempotency keys, use them
Architecture suggestion
- API service: creates email jobs
- Queue: SQS, RabbitMQ, Kafka, Redis/BullMQ, etc.
- Worker(s): perform send logic
- DB: track status, attempts, provider used, timestamps
- Metrics/logging: monitor failures, latency, bounce rates
Data model example
Store fields like:
idrecipientsubjectbodystatus(pending,sent,failed)attempt_countproviderlast_errornext_retry_atcreated_at,updated_at
Best practices
- Separate submission from delivery
- Keep retries in the worker, not the API request
- Add circuit breaker logic if a provider is down
- Monitor bounce/complaint webhooks
- Handle rate limiting from providers
- Have a dead-letter queue for exhausted jobs
Simple logic
Pseudo-flow:
for each email job:
try primary provider
if transient error:
retry with backoff
if retries exhausted:
try backup provider
if success:
mark sent
else:
mark failed
If you want something production-ready
A solid stack could be:
- Queue: SQS or RabbitMQ
- Worker: Node.js/Python/Go
- Primary: SES
- Backup: SendGrid or Mailgun
- Observability: CloudWatch / Prometheus / Datadog
If you want, I can also give you:
- a reference architecture diagram
- sample code in Node.js, Python, or Go
- a retry/failover policy design with exact error classification