Prompt
How do I handle retries and fallback models with an LLM API?
Latest observation
A good pattern is to treat retries and fallback models as client-side orchestration, not something you leave to chance.
1) Retry only for transient failures
Retry when the failure is likely temporary, for example:
- HTTP
429rate limit - HTTP
500,502,503,504 - network timeouts / connection resets
Don’t blindly retry for:
- invalid API key
- bad request / malformed prompt
- context too long
- content policy rejections
2) Use exponential backoff with jitter
Avoid retrying immediately in a tight loop. Instead:
- wait a short time
- then increase delay on each retry
- add jitter so many clients don’t retry at once
Typical pattern:
- attempt 1: no wait
- attempt 2: 200–500 ms
- attempt 3: 1–2 s
- attempt 4: 4–8 s
Cap the maximum delay and the total retry time.
3) Keep retries bounded
Set clear limits:
- max attempts: e.g. 3–5
- max total elapsed time
- idempotency or deduplication if your workflow can cause duplicates
For generation requests, retries are usually safe, but if the call triggers side effects downstream, ensure the full workflow is idempotent.
4) Add fallback models by priority
If your primary model fails after retries, fall back to a cheaper or more available model:
Example order:
gpt-4.1gpt-4.1-minigpt-4o-mini
Use fallback for:
- repeated transient failures
- overload
- cost control
- latency constraints
Be careful: fallback models may have different quality, formatting, or tool-use behavior.
5) Normalize prompts across models
To make fallback safe:
- keep system instructions consistent
- avoid model-specific prompt tricks
- request structured output when possible
- validate output before using it
If the fallback model returns malformed output, either:
- repair/parse it
- or escalate to a human / alternate workflow
6) Separate failure types
A practical policy:
- Retry same model for transient errors
- Fallback to another model after retry budget is exhausted
- Fail fast on permanent errors
7) Log enough to debug
Record:
- model used
- attempt count
- error type/status code
- latency
- whether fallback was used
- final response quality metrics if you have them
This helps you tune retry thresholds and fallback ordering.
8) Example strategy
A simple policy:
- Try primary model up to 3 times
- Retry on 429, 5xx, timeouts
- Backoff: 0.5s, 1s, 2s with jitter
- If still failing, switch to fallback model and try 2 times
- If all fail, return a controlled error
9) Python-style pseudocode
import random
import time
RETRYABLE_STATUSES = {429, 500, 502, 503, 504}
def call_model(model, prompt):
# replace with actual API call
pass
def should_retry(error):
status = getattr(error, "status_code", None)
return status in RETRYABLE_STATUSES or isinstance(error, TimeoutError)
def run_with_fallback(prompt):
models = [
("primary-model", 3),
("fallback-model", 2),
]
for model, max_attempts in models:
for attempt in range(1, max_attempts + 1):
try:
return call_model(model, prompt)
except Exception as e:
if not should_retry(e):
raise
if attempt == max_attempts:
break
base_delay = 0.5 * (2 ** (attempt - 1))
jitter = random.uniform(0, base_delay * 0.2)
time.sleep(base_delay + jitter)
raise RuntimeError("All models failed")
10) Best practice summary
- Retry transient failures only
- Use exponential backoff + jitter
- Cap retries
- Fall back to another model after retry exhaustion
- Validate outputs from fallback models
- Log and monitor everything
If you want, I can also give you:
- a production-ready Python implementation
- a Node.js example
- or a retry/fallback policy template for your specific API stack