Prompt
How do I add retries and failover to API requests?
Latest observation
To add retries and failover to API requests, you want a pattern that handles:
- Transient failures with retry
- Endpoint outages with failover to another base URL/region/provider
- Safe retry behavior so you don’t duplicate side effects
1) Retry only when it makes sense
Retry on:
- timeouts
- connection errors
- DNS issues
- 429 rate limits
- 5xx server errors
Usually do not retry on:
- most 4xx errors like 400, 401, 403, 404
- requests that aren’t idempotent unless you have idempotency keys
Use:
- a small max retry count, e.g. 2–5
- exponential backoff with jitter
- an overall timeout
Example backoff:
- 1s
- 2s
- 4s
- plus random jitter
2) Add failover endpoints
Maintain a list of endpoints, such as:
- primary API region
- secondary API region
- backup provider
On repeated failures, switch to the next endpoint.
A common approach:
- Try primary endpoint
- Retry a few times with backoff
- If still failing, move to secondary
- Optionally “circuit break” the bad endpoint for a cooldown period
3) Make requests idempotent
For operations like payments, order creation, or ticket submission:
- use idempotency keys
- or ensure the API supports safe retries
This prevents duplicate actions if a request succeeds but the response is lost.
4) Use timeouts everywhere
Set:
- connect timeout
- read timeout
- overall request timeout
Without timeouts, retries can hang and make failures worse.
5) Add circuit breaking
If an endpoint is failing repeatedly:
- stop sending traffic to it temporarily
- let it recover
- then probe it again
This prevents thundering herds and wasted retries.
Example: Python requests with retries
Using urllib3 retry support:
import requests
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry
session = requests.Session()
retry = Retry(
total=3,
backoff_factor=1,
status_forcelist=[429, 500, 502, 503, 504],
allowed_methods=["GET", "POST", "PUT", "DELETE", "PATCH"],
raise_on_status=False,
)
adapter = HTTPAdapter(max_retries=retry)
session.mount("https://", adapter)
session.mount("http://", adapter)
response = session.get("https://api.example.com/data", timeout=5)
print(response.status_code, response.text)
Example: simple manual failover in Python
import requests
import time
import random
ENDPOINTS = [
"https://primary.api.example.com",
"https://secondary.api.example.com",
]
def call_api(path, method="GET", **kwargs):
last_error = None
for base_url in ENDPOINTS:
url = base_url + path
for attempt in range(3):
try:
resp = requests.request(method, url, timeout=5, **kwargs)
if resp.status_code in (429, 500, 502, 503, 504):
raise requests.HTTPError(f"Retryable status: {resp.status_code}", response=resp)
return resp
except (requests.Timeout, requests.ConnectionError, requests.HTTPError) as e:
last_error = e
sleep_time = (2 ** attempt) + random.uniform(0, 0.5)
time.sleep(sleep_time)
raise last_error
resp = call_api("/data")
print(resp.status_code, resp.text)
6) Recommended best practices
- Prefer retries in a shared HTTP client wrapper
- Log retry count, endpoint, and error type
- Add metrics/alerts for failover events
- Use jitter to avoid synchronized retries
- Don’t retry too aggressively
- Respect
Retry-Afterheaders for 429/503 if provided
7) If you’re using a language/framework
I can give you a ready-made example for:
- Python
- JavaScript/Node.js
- Java
- Go
- C#
- Ruby
If you want, tell me your language and whether your requests are read-only or side-effecting (like creating payments/orders), and I’ll show the safest implementation.