Prompt
How do I set up failover for critical application emails?
Latest observation
To set up failover for critical application emails, you want a design that keeps mail delivery working even if your primary email service, SMTP relay, DNS provider, or app environment has issues.
A practical failover pattern
1. Use a primary and secondary mail delivery path
Set up:
- Primary SMTP provider: normal delivery path
- Secondary SMTP provider: fallback if primary fails
Your application should be able to:
- Try the primary first
- Detect failure/timeouts
- Automatically retry through the secondary provider
This is usually done in one of three ways:
- App-level failover logic
- Mail gateway / relay service with failover
- Messaging queue + delivery worker with provider switching
2. Put emails into a queue
For critical emails, don’t send them directly from the request thread.
Use a queue so that:
- App requests aren’t blocked
- Temporary provider outages don’t lose messages
- A worker can retry delivery later
Example flow:
- Application creates email event
- Message is stored in a queue/database
- Worker sends via primary SMTP
- If failed, worker retries or switches to secondary SMTP
- If all fail, alert operations and keep retrying
3. Define failover rules
Common rules:
- Timeout: if SMTP doesn’t respond within X seconds, fail over
- Connection error: switch immediately to secondary
- 5xx SMTP errors: treat as provider/system failure and fail over
- 4xx SMTP errors: usually retry primary first, since these may be temporary
Good failover logic distinguishes between:
- Transient failures → retry
- Service outages → fail over
- Invalid recipient/content issues → do not fail over blindly
4. Add health checks
Monitor each email provider’s health:
- Successful send rate
- Latency
- Error rate
- Bounce and complaint rates
If primary is degraded, route new sends to secondary until recovered.
A simple approach:
- Primary is default
- If 3 consecutive send failures or high error rate in 5 minutes, mark primary unhealthy
- After a cooldown, test primary again
5. Use DNS/MX failover only if you run your own mail servers
If you operate your own mail infrastructure, set up:
- Primary MX
- Backup MX with higher priority number
- Proper SPF, DKIM, and DMARC alignment
But note:
- MX failover is for receiving mail, not directly for app-originated outbound delivery
- For sending, you still need SMTP relay redundancy
6. Make sure messages are idempotent
When retrying or failing over, you risk duplicate emails.
Prevent duplicates by:
- Assigning a unique message ID
- Storing send status
- Checking whether an email was already accepted by a provider before retrying
- Designing application emails to tolerate duplicate delivery when necessary
7. Monitor and alert
Set alerts for:
- Primary provider outage
- Secondary provider usage spikes
- Queue backlog growth
- Delivery failure rate above threshold
- Bounce/complaint anomalies
Also keep a dashboard for:
- Queue depth
- Send success rate
- Provider latency
- Retry counts
8. Test failover regularly
Don’t assume it works until tested.
Run tests for:
- SMTP timeout simulation
- Provider API failure
- DNS resolution failure
- Credential expiration
- Network interruption
Verify:
- Messages move to secondary
- No lost emails
- No duplicate bursts
- Alerting works
Recommended architecture
For most critical applications:
- App → Queue → Email worker
- Worker sends through Primary SMTP provider
- On failure, worker switches to Secondary provider
- If both fail, keep retrying with backoff
- Alerts fire when backlog or failure rate rises
Example failover policy
- Try provider A
- If timeout > 5 seconds or 5xx error:
- mark provider A unhealthy
- send via provider B
- Retry failed messages every 1, 5, 15, 60 minutes
- After 10 minutes of stability, probe provider A again
Common mistakes to avoid
- Sending directly from web requests without a queue
- Using only one SMTP provider
- Treating all SMTP errors the same
- Not alerting on backlog or failures
- Not testing failover
- Causing duplicate emails without tracking
If you want, I can also give you:
- a cloud-provider-specific setup (AWS SES, SendGrid, Mailgun, etc.),
- a reference architecture diagram, or
- sample code for automatic SMTP failover.