Prompt
What should I use for multi-model fallback and retries?
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
If you want multi-model fallback plus retries, the best pattern is usually:
Use retries for transient failures
Retry the same model when the problem is likely temporary:
- timeouts
- 429 rate limits
- 5xx server errors
- network errors
Good retry policy:
- exponential backoff
- jitter
- small max attempts, usually 2–3
- don’t retry on bad requests / prompt errors
Use fallback for model-level failure or quality issues
Switch to another model when:
- the primary model is unavailable for too long
- the request repeatedly times out
- you hit persistent rate limits
- you need a cheaper/faster backup
- you want a stronger model for hard cases
Typical fallback order:
- fast/cheap model first
- stronger model second
- most reliable model last
Best practice: combine both
A common setup is:
- Try primary model
- Retry it a couple times on transient errors
- If still failing, move to fallback model
- Log which model answered, latency, and error type
If you’re asking what to use in code
Use an orchestration layer that supports:
- retries with backoff
- fallback routing
- timeout handling
- observability
If you’re building this yourself, a simple policy engine is often enough.
Rule of thumb
- Retries = same model, same request, temporary failures
- Fallbacks = different model, after repeated failure or according to policy
If you want, I can suggest:
- a recommended policy for production,
- a Python example, or
- how to do this with LangChain / LiteLLM / OpenAI-style client code.