Prompt
LiteLLM retries cache rate limits
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
If you want LiteLLM to handle retries, cache, and rate limits, there are a few different layers involved:
1) Retries
LiteLLM can retry failed requests, but the exact behavior depends on how you’re calling it.
Typical approach:
- configure the
num_retries/ retry settings in the LiteLLM client or proxy - retry on transient errors like
429,500, timeouts
Example pattern:
import litellm
response = litellm.completion(
model="openai/gpt-4o-mini",
messages=[{"role": "user", "content": "Hello"}],
num_retries=3,
)
2) Cache
LiteLLM supports caching so repeated identical requests can be served without hitting the provider again.
Common options:
- in-memory cache
- Redis cache
- prompt/response caching via proxy or SDK settings
Example concept:
from litellm import completion
from litellm.caching import Cache
litellm.cache = Cache(type="redis")
Caching is useful to reduce:
- latency
- token usage
- provider quota consumption
3) Rate limits
There are two kinds of rate limits:
- Provider-side limits (OpenAI, Anthropic, etc.)
- Your own app-side limits
LiteLLM can help with provider throttling by:
- retrying on
429 - routing across fallbacks
- using proxy-level rate limiting / budgets
If you’re using the LiteLLM Proxy, you can also configure:
- per-model RPM/TPM limits
- budgets
- fallback models
Practical recommendation
A good setup is:
- Enable cache for repeated prompts
- Set retries for transient failures
- Use rate limiting / fallback models to avoid exhausting one provider
- Log 429s so you can see whether you need more aggressive throttling
Example strategy
- cache identical requests for 1 hour
- retry up to 3 times with exponential backoff
- on 429, switch to a fallback model/provider
- optionally enforce local RPM limits before sending requests
If you want, I can give you:
- a LiteLLM Python SDK example, or
- a LiteLLM Proxy config showing retries + cache + rate limits together.