Prompt
How do I set fallback rules for LLM requests when a provider fails?
Latest observation
A good way to handle LLM provider failures is to define a fallback policy with clear priorities, retries, and routing rules. Here’s a practical pattern you can use.
1) Decide what counts as a failure
Common failure types:
- Timeout
- Rate limit / 429
- 5xx provider errors
- Network errors
- Malformed response / schema violation
- Safety refusal (depending on your use case)
Not every failure should trigger the same fallback. For example:
- Retry the same provider for transient errors
- Switch providers for persistent or capacity-related errors
- Return a cached or lower-cost model for non-critical requests
2) Define a fallback chain
A simple chain might look like:
- Try primary provider/model
- If timeout/5xx/429, retry once with backoff
- If still failing, try secondary provider/model
- If still failing, try a cheaper or smaller model
- If everything fails, return a graceful degradation message or queue for later
Example policy:
primary: openai/gpt-4.1
fallback_1: anthropic/claude-3.5-sonnet
fallback_2: openai/gpt-4o-mini
final: cached answer / apology / manual review
3) Use error-specific rules
Map errors to actions:
- 429 / rate limit → wait + retry, then fallback
- Timeout → retry with shorter timeout, then fallback
- 5xx → retry once, then fallback
- Invalid output → retry same model with stricter prompt or schema, then fallback
- Auth / permission errors → do not retry; fail fast and alert
4) Keep request compatibility in mind
Fallback models may differ in:
- context window
- tool-calling support
- structured output reliability
- temperature/parameter support
So your fallback rule should also specify:
- whether to truncate context
- whether to disable tools
- whether to relax schema validation
- whether to downgrade to a simpler prompt
5) Add circuit breakers
If a provider is failing repeatedly, stop sending traffic to it for a cooldown period.
Typical behavior:
- After N failures in M minutes, open the circuit
- Route all requests to fallbacks
- Periodically probe the primary provider
- Close the circuit when healthy again
This avoids hammering a broken provider.
6) Prefer deterministic routing
Don’t make fallback decisions ad hoc in code scattered everywhere. Use a central policy:
- ordered provider list
- per-error retry counts
- cooldown durations
- request class priorities
For example:
{
"primary": "openai:gpt-4.1",
"fallbacks": [
{"provider": "anthropic:claude-3.5-sonnet", "on": ["timeout", "rate_limit", "server_error"]},
{"provider": "openai:gpt-4o-mini", "on": ["timeout", "rate_limit", "server_error", "schema_error"]}
],
"retry": {
"timeout": 1,
"rate_limit": 2,
"server_error": 1
},
"circuit_breaker": {
"failure_threshold": 5,
"window_seconds": 60,
"cooldown_seconds": 120
}
}
7) Log and measure fallback usage
Track:
- fallback rate
- reason for fallback
- latency added by retries
- success rate per provider
- user impact
This helps you tune the policy and identify flaky providers.
8) Have graceful degradation
If no provider works:
- return a partial answer
- ask the user to retry
- save the request for later processing
- fall back to a template response
- use retrieval/cached content if available
9) If you’re using a framework or gateway
Many LLM gateways/proxies support this directly with:
- provider failover
- load balancing
- weighted routing
- health checks
- retries and circuit breakers
If you build it yourself, the main logic is:
- call provider
- classify error
- retry if appropriate
- move to next provider in policy
- fail gracefully
Example pseudo-code
providers = [
PrimaryProvider(),
SecondaryProvider(),
CheapFallbackProvider(),
]
for provider in providers:
try:
return provider.generate(prompt, timeout=10)
except TransientError:
continue
except InvalidSchemaError:
# maybe retry once with stricter prompt here
continue
raise RuntimeError("All LLM providers failed")
If you want, I can also give you:
- a production-grade fallback policy design
- an example in Python/TypeScript
- a setup for OpenAI + Anthropic + local model failover