Prompt
How do I add retries, timeouts, and failover at the API gateway layer?
Latest observation
To add retries, timeouts, and failover at the API gateway layer, you typically configure them in the gateway itself or in the gateway’s upstream/service policies. The exact syntax depends on the gateway product, but the design pattern is the same.
1) Timeouts
Set limits for how long the gateway will wait for:
- Connection timeout: time to establish a connection to the backend
- Read/response timeout: time to wait for the backend response
- Overall request timeout: total time allowed end-to-end
Best practice
- Keep gateway timeouts slightly lower than client timeouts.
- Set upstream timeouts based on backend SLA and expected latency.
- Use different timeouts per route if some endpoints are slower.
2) Retries
Configure retries for transient failures only, such as:
- connection errors
- DNS resolution failures
- 502/503/504 responses
- timeouts
Best practice
- Retry only idempotent methods by default:
GET,HEAD,PUT,DELETE(depending on your API semantics)
- Avoid retrying
POSTunless you have:- idempotency keys
- deduplication on the backend
- Use:
- max retry attempts
- retry conditions
- backoff strategy with jitter
Example retry policy
- 2 retries
- exponential backoff: 100ms, 200ms
- retry on 502, 503, 504, and connect/read timeout
3) Failover
Failover means routing traffic to an alternate backend when the primary fails.
Common approaches:
- Active-passive: primary backend, fallback backend if primary is unhealthy
- Active-active: distribute across multiple backends; remove unhealthy ones
- Region failover: route to another region during outages
Best practice
- Use health checks to determine upstream health.
- Fail over only on specific failure types.
- Prefer local retry first, then failover to secondary.
- Avoid thrashing by using circuit breakers and outlier detection.
4) Circuit breakers and outlier detection
These are often paired with retries/failover:
- Circuit breaker: stop sending requests to a failing upstream after too many errors
- Outlier detection: eject unhealthy instances from load balancing
This prevents retry storms and cascading failures.
5) A practical policy pattern
For each route:
- Set connect/read timeouts
- Retry 1–3 times on transient errors
- Use exponential backoff + jitter
- Fail over to secondary upstream if primary is unhealthy
- Protect with circuit breaker/outlier detection
- Monitor latency, error rate, and retry volume
6) Example configurations
Kong
- Use plugins such as:
request-terminationfor fallback responses- upstream settings for timeouts and retries
- health checks / load balancing via upstream entities
NGINX / NGINX Plus
Typical directives:
proxy_connect_timeoutproxy_read_timeoutproxy_send_timeoutproxy_next_upstreamproxy_next_upstream_tries- upstream health checks / load balancing
Envoy
Typical settings:
timeoutretry_policyper_try_timeoutretry_on- upstream health checks
- outlier detection
AWS API Gateway
API Gateway itself has limited retry/failover control compared with proxies; failover is often implemented with:
- Lambda, ALB, Route 53, or CloudFront
- multi-region backend design
- custom integration logic
7) Operational tips
- Instrument:
- retry count
- timeout count
- upstream error rates
- failover events
- Alert on:
- rising retries
- latency near timeout thresholds
- circuit breaker open events
- Test:
- backend latency injection
- 5xx failures
- regional outages
- DNS failures
8) Important cautions
- Retries can amplify outages if misused.
- Timeouts that are too short can cause false failures.
- Failover can create consistency issues if stateful operations are routed elsewhere.
- Avoid retrying non-idempotent operations without safeguards.
If you tell me which API gateway you’re using (Kong, NGINX, Envoy, Apigee, AWS API Gateway, Azure API Management, etc.), I can give you the exact configuration example.