Prompt
How do I set up fallback between multiple LLM providers if one is slow or down?
Latest observation
A good fallback setup for multiple LLM providers usually has 4 parts:
- A common interface
- Health/latency tracking
- A routing policy
- Retries + timeouts + circuit breakers
1) Normalize all providers behind one interface
Wrap each provider so your app calls them the same way:
class LLMProvider:
def generate(self, prompt, **kwargs):
raise NotImplementedError
Then implement one adapter per provider:
class OpenAIProvider(LLMProvider):
def generate(self, prompt, **kwargs):
...
class AnthropicProvider(LLMProvider):
def generate(self, prompt, **kwargs):
...
class GeminiProvider(LLMProvider):
def generate(self, prompt, **kwargs):
...
This makes failover logic provider-agnostic.
2) Set strict timeouts and retries
Don’t wait forever on a slow provider.
- Connect timeout: e.g. 2–5s
- Request timeout: e.g. 15–30s depending on use case
- Retries: small number, only for transient errors
- Backoff: exponential with jitter
Important:
- Retry timeouts, 429s, 5xx
- Don’t retry most 4xx errors like bad auth or invalid input
Example policy:
- Try provider A
- If A times out or fails transiently, try B
- If B fails, try C
- Stop after N providers or return an error
3) Track provider health
Maintain live stats per provider:
- success rate
- p50 / p95 latency
- timeout rate
- error rate
- last failure time
- circuit breaker state
Use this to decide which provider to try first.
A simple rule:
- Prefer the fastest healthy provider
- Demote providers with high error/timeout rates
- Temporarily disable a provider after repeated failures
4) Use a circuit breaker
If a provider is repeatedly failing, stop sending traffic to it for a short cooldown.
Typical states:
- Closed: normal
- Open: do not send requests
- Half-open: test with a small number of requests
This prevents wasting time on a broken provider.
Recommended request flow
- Build a ranked provider list based on:
- health
- latency
- cost
- model capability
- Send request to the top provider with a timeout
- If it fails or is too slow:
- log the failure
- mark provider health
- move to next provider
- Return first successful response
Example fallback pseudocode
def generate_with_fallback(prompt, providers):
last_error = None
for provider in providers:
try:
return provider.generate(prompt, timeout=20)
except TimeoutError as e:
last_error = e
continue
except TransientProviderError as e:
last_error = e
continue
except PermanentProviderError:
# Usually don't retry on permanent config/auth errors
raise
raise RuntimeError(f"All providers failed: {last_error}")
Better: rank by health before trying
def rank_providers(providers, health):
return sorted(
providers,
key=lambda p: (
not health[p.name].is_healthy(),
health[p.name].p95_latency,
health[p.name].error_rate
)
)
Then try them in ranked order.
Best practices
Fail open vs fail closed
Decide what to do if all providers fail:
- Fail open: return partial response, cached answer, or simplified mode
- Fail closed: return an error immediately
Separate by task
Use different fallback policies for:
- chat
- embeddings
- tool/function calling
- long-context summarization
A provider that is good for chat may not be good for tool calling.
Keep prompts compatible
Different providers have different formatting requirements:
- system messages
- tool schemas
- token limits
- sampling params
Normalize prompt construction in one place.
Watch out for response differences
Fallback may change:
- style
- safety behavior
- function-calling format
- JSON strictness
If you need structured output, enforce validation and re-ask if needed.
If you want production-grade reliability
Use:
- timeouts
- retries
- circuit breaker
- health scoring
- request hedging for very latency-sensitive requests
Hedging
Instead of waiting for a slow provider to fail, you can:
- send to provider A
- after a short delay, send to B as backup
- take the first valid response This reduces tail latency but costs more.
Simple architecture
Client → Router → Provider Adapters → Monitoring
- Router picks provider
- Adapters handle API differences
- Monitoring feeds latency/error metrics back into routing
Minimal practical setup
If you want the simplest robust version:
- Use 2–3 providers
- Add a 10–20s request timeout
- Retry once on transient errors
- If the request fails, try the next provider
- Track success/latency in Redis or a DB
- Add a circuit breaker after 3–5 consecutive failures
If you want, I can give you:
- a Python implementation of an LLM fallback router, or
- a LangChain / LiteLLM / OpenRouter-based setup.