Prompt
How do I do failover routing when an API backend is down?
Latest observation
Failover routing for an API backend usually means: try the primary backend first, and if it’s unhealthy or unreachable, send traffic to a standby backend automatically.
Here are the common ways to do it, from simplest to most robust.
1) Put a load balancer/API gateway in front
This is the standard approach.
How it works
- Clients call a single endpoint
- A load balancer or API gateway routes requests to backend instances
- Health checks detect when an instance or region is unhealthy
- Traffic is shifted to healthy targets
Tools
- AWS: ALB / NLB, API Gateway + Lambda/ALB, Route 53 health checks
- GCP: Cloud Load Balancing, API Gateway, Cloud DNS health checks
- Azure: Front Door, Application Gateway, Traffic Manager
- Kubernetes: Ingress controller + service health checks, or service mesh
Best for
- Multiple instances in one region
- Active-active or active-passive across zones/regions
2) Use active-passive failover
One backend is primary, another is standby.
Flow
- Send traffic to primary backend
- Monitor health
- If primary fails health checks, switch traffic to secondary
- Switch back only after primary is healthy and stable
Pros
- Simple
- Easy to reason about
- Lower risk of split-brain issues
Cons
- Secondary may be idle and less tested
- Failover can take seconds to minutes depending on detection
Important
Use a failback policy carefully. Don’t switch back too quickly or you may flap between backends.
3) Use DNS-based failover
Route traffic with DNS records that change when health checks fail.
Example
api.example.compoints to primary- If primary is unhealthy, DNS returns secondary IP/hostname
Caveat
DNS failover can be slow because of:
- TTL caching
- Client-side DNS caching
- Resolver caching
Good for
- Region-level failover
- Simpler setups
Not ideal for
- Fast failover requirements
4) Retry with circuit breaker in the client or gateway
If the backend fails, don’t keep hammering it.
Pattern
- Request fails
- Retry once or twice with backoff
- If failures exceed a threshold, open a circuit breaker
- Route to secondary backend while primary is unhealthy
Important
Use retries carefully:
- Retry only idempotent operations unless you have idempotency keys
- Add exponential backoff + jitter
- Set short timeouts
Typical rule
- Timeout: 1–3 seconds for API calls
- Retries: 1–2 max
- Circuit breaker opens after a failure threshold
5) Multi-region active-active
Both backends serve traffic all the time.
How it works
- Requests go to the nearest or healthiest region
- Global traffic management handles routing
- If a region goes down, traffic continues elsewhere
Pros
- Best availability
- Better disaster tolerance
Cons
- Harder data consistency
- More complex architecture
- Requires replication strategy
Good for
- Mission-critical APIs
- Global apps
Recommended pattern
For most APIs, a good setup is:
- Load balancer or API gateway
- Health checks
- Timeouts
- Retries with backoff
- Circuit breaker
- Secondary backend/region
- Monitoring + alerting
Example architecture
Single region
- API Gateway
- Primary app instances in AZ A and AZ B
- Health checks
- Auto-scaling group
- If all primary instances fail, route to standby service
Multi-region
- Global load balancer
- Region A primary
- Region B standby
- Health checks monitor regional endpoints
- Automatic DNS or traffic-policy failover
Things to watch out for
1) Data consistency
If your API writes data, failover can cause issues unless both backends share or replicate state properly.
Questions to answer:
- Are reads/writes going to the same database?
- Is the DB replicated?
- Is replication synchronous or asynchronous?
2) Idempotency
If a request is retried during failover, avoid duplicate side effects. Use:
- Idempotency keys
- Request deduplication
- Transaction IDs
3) Statefulness
Sticky sessions can break failover if session state lives only on one backend. Prefer:
- Stateless API servers
- Shared session store like Redis
- Token-based auth
4) Failover detection speed
Health checks that are too aggressive can cause false positives. Balance:
- detection speed
- false failure rate
- recovery stability
Example: practical failover logic
If you’re implementing it in a gateway or middleware:
- Call primary backend with 2s timeout
- If timeout/5xx:
- increment failure counter
- If failure threshold exceeded:
- mark primary unhealthy for 30s
- route to secondary
- Periodically probe primary to restore it
If you want a concrete implementation
Tell me what stack you’re using, and I can give you a specific setup, for example:
- AWS ALB + Route 53 failover
- Nginx upstream failover
- Kubernetes ingress
- Node.js / Python client failover logic
- Cloudflare / Fastly / API gateway failover
If you want, I can also show:
- a sample Nginx config
- a sample Node.js failover client
- an AWS architecture
- a Kubernetes solution