Prompt
How do I set up retries and fallbacks for an LLM API?
Latest observation
Oct 8, 2026
OpenAI APIWeb search: off
A good retry/fallback setup for an LLM API usually has three layers:
- Retry transient failures
- Fallback to a cheaper or more reliable model/provider
- Degrade gracefully when all else fails
1) Retry only the failures that are worth retrying
Retry:
- timeouts
- rate limits (
429) - transient server errors (
500,502,503,504) - network errors
Do not retry:
- invalid prompts
- auth errors
- bad request/schema errors
- content policy failures
Use:
- exponential backoff
- jitter to avoid thundering herds
- a max retry count
- an overall timeout budget
Example retry policy
- Try up to 3 times
- Backoff:
0.5s,1s,2s - Add random jitter ±20%
- Stop if total elapsed time exceeds 10s
2) Add fallback paths
Common fallback order:
Model fallback
- Primary: best-quality model
- Fallback 1: cheaper/smaller model from same provider
- Fallback 2: another provider
- Fallback 3: rules-based or cached response
Request simplification fallback
If a request is too expensive or large:
- shorten context
- reduce max tokens
- disable optional tools/function calling
- switch to a summarization or extraction prompt
Provider fallback
If one provider is down:
- route to another provider with a similar model
- normalize prompts/responses across providers
3) Define your fallback decision logic
Use different policies for different failure types:
- 429 / rate limit → retry with longer backoff, maybe switch to a lower-priority model
- 5xx / timeout → retry same provider first, then fallback model/provider
- context too long → automatically trim/summarize, don’t retry unchanged
- invalid response format → one retry with a stricter prompt/schema, then fallback
4) Protect your app with guardrails
Add:
- circuit breaker: stop calling a failing provider for a cooldown period
- request deduplication: avoid repeated identical retries
- idempotency keys where supported
- observability: log attempt count, latency, error type, fallback used
- budget limits: cap cost per request
5) A simple implementation pattern
Pseudo-logic:
def call_llm(request):
attempts = [
{"model": "best-model", "provider": "A"},
{"model": "smaller-model", "provider": "A"},
{"model": "backup-model", "provider": "B"},
]
for attempt in attempts:
for retry in range(3):
try:
return llm_api_call(request, **attempt)
except RetryableError as e:
sleep(backoff_with_jitter(retry))
continue
except NonRetryableError:
break
return fallback_response()
6) Practical recommendations
- Keep retries short and bounded
- Don’t retry blindly on every error
- Fall back to a known-safe response if the LLM is unavailable
- Prefer structured outputs and validate them before accepting them
- Log enough data to know whether retries are helping or just increasing cost
7) If you want a production-ready setup
Typical stack:
- retry middleware with exponential backoff
- circuit breaker
- multi-provider router
- prompt/version registry
- response validator
- fallback cache or templated response
If you want, I can show you:
- a Python example using
tenacity - a Node.js example
- a multi-provider fallback architecture
- or a production checklist for LLM retries/failovers
Brands