Prompt
What should I use for retries, caching, and failover on AI calls?
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
A good default is to treat AI calls like flaky external APIs and put them behind a small resilience layer with:
- Retries
- Caching
- Failover / fallback routing
1) Retries
Use retries for transient failures only:
- network timeouts
- 429 rate limits
- 5xx provider errors
- connection resets
Best practices:
- Exponential backoff with jitter
- Limit attempts to 2–4 max
- Set timeouts on every request
- Do not blindly retry on bad prompts, auth errors, or validation errors
- For streaming responses, only retry if your app can safely reissue the full request
Good tooling:
- Tenacity (Python)
- Polly (.NET)
- resilience4j (Java)
- p-retry / custom wrapper (Node)
- Or your cloud’s retry middleware if it supports backoff + jitter
2) Caching
Cache when the result is likely to be reused and can tolerate being slightly stale.
Great cache targets:
- embeddings
- classification outputs
- prompt templates with static context
- deterministic summarizations
- repeated Q&A with the exact same input
- tool results fetched from external systems
Avoid caching or be careful with:
- highly personalized outputs
- fresh data requests
- anything sensitive unless encryption/TTL controls are strong
Cache types:
- Exact-match cache for identical prompts/params
- Semantic cache for “similar” prompts
- Response cache with TTL
- Tool/result cache for upstream API calls
Good tooling:
- Redis for shared low-latency cache
- Memcached for simple ephemeral caching
- SQLite / local LRU for single-instance apps
- Vector DB or semantic cache layer for similarity-based reuse
3) Failover
Use failover when availability matters more than sticking to one model/provider.
Options:
- Same model, different region
- Primary provider, secondary provider
- Primary model, cheaper/smaller backup model
- Graceful degradation: shorter answer, lower token limit, non-streaming, or cached fallback
Best practices:
- Define a provider priority order
- Use a circuit breaker so you stop hammering a failing provider
- Monitor quality differences between providers
- Normalize request/response formats in one abstraction layer
- Keep prompts compatible across models
Good tooling:
- Circuit breaker libraries like resilience4j / Polly
- Multi-provider routers such as:
- LiteLLM
- custom gateway/router
- API gateway with retry + routing logic
Practical recommendation
If you want a simple, solid stack:
- Retries: exponential backoff + jitter, max 3 attempts
- Cache: Redis with TTL for exact matches, plus optional semantic cache
- Failover: circuit breaker + secondary model/provider
- Observability: log model, latency, retries, cache hit rate, and error type
Rule of thumb
- If the call is expensive and repeatable → cache it
- If the failure is temporary → retry it
- If the provider is down or degraded → fail over
If you want, I can suggest a concrete stack for Python, Node.js, or AWS/GCP/Azure.