Prompt
LLM API with retries and fallbacks
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
Here’s a practical pattern for an LLM API wrapper with retries and fallbacks.
What it should do
- Retry on transient failures:
- timeouts
- rate limits
- 5xx errors
- network errors
- Fallback to a cheaper/smaller model or alternate provider if the primary keeps failing
- Use exponential backoff + jitter
- Cap total latency with a timeout budget
- Log/metric each attempt for observability
Example design
Flow
- Try primary model
- If it fails, retry
ntimes with backoff - If still failing, try fallback model(s)
- If all fail, raise a clean error
Python example
import random
import time
from typing import List, Callable, Any
class LLMError(Exception):
pass
class RateLimitError(LLMError):
pass
class TimeoutError(LLMError):
pass
class TemporaryServerError(LLMError):
pass
def is_retriable(err: Exception) -> bool:
return isinstance(err, (RateLimitError, TimeoutError, TemporaryServerError))
def exponential_backoff(attempt: int, base: float = 0.5, cap: float = 8.0) -> float:
delay = min(cap, base * (2 ** attempt))
jitter = random.uniform(0, delay * 0.2)
return delay + jitter
def call_with_retries(
fn: Callable[[], Any],
max_retries: int = 3,
) -> Any:
last_err = None
for attempt in range(max_retries + 1):
try:
return fn()
except Exception as err:
last_err = err
if attempt == max_retries or not is_retriable(err):
raise
time.sleep(exponential_backoff(attempt))
raise last_err
def llm_request(prompt: str, model: str) -> str:
"""
Replace this with your provider call, e.g. OpenAI/Anthropic/etc.
"""
# Simulated behavior:
if model == "primary":
raise TemporaryServerError("Primary model unavailable")
return f"[{model}] response to: {prompt}"
def generate_with_fallbacks(prompt: str, models: List[str]) -> str:
last_err = None
for model in models:
try:
return call_with_retries(lambda: llm_request(prompt, model), max_retries=3)
except Exception as err:
last_err = err
print(f"Model {model} failed: {err}")
raise LLMError(f"All models failed. Last error: {last_err}")
if __name__ == "__main__":
result = generate_with_fallbacks(
"Write a short product description for headphones.",
models=["primary", "fallback-small", "fallback-cheap"]
)
print(result)
Recommended production features
1. Retry only safe errors
Don’t retry:
- invalid auth
- bad request / prompt formatting
- context length exceeded
- content policy violations
Retry:
- 429
- 500, 502, 503, 504
- connection resets
- timeouts
2. Add timeout budget
Example:
- total request budget: 15s
- per attempt: 5s
- retries/backoff must fit inside budget
3. Model fallback strategy
A common chain:
gpt-4.1for best qualitygpt-4.1-minifor cheaper fallbackgpt-4o-minifor last-resort fast fallback
Or across providers:
- OpenAI → Anthropic → local model
4. Circuit breaker
If a provider is failing repeatedly, stop calling it for a short window to avoid hammering it.
5. Structured logging
Log:
- request id
- model name
- attempt number
- latency
- error type
- fallback used
If you want, I can also give you:
- a real OpenAI/Anthropic implementation
- a TypeScript version
- a version with async/await
- a version using Tenacity for retries