Prompt
How do I add caching and retries to LLM requests without rewriting my app?
Latest observation
The easiest way is to put a wrapper/proxy layer between your app and the LLM provider, so your app keeps calling the same API shape while the wrapper handles caching, retries, rate limits, logging, and fallback behavior.
Option 1: Use a drop-in client wrapper
If your app already uses a standard SDK, wrap the call with a small helper:
- Retries for transient failures: timeouts, 429s, 5xx
- Cache for repeated prompts
- Backoff + jitter to avoid retry storms
Example pattern
from tenacity import retry, stop_after_attempt, wait_exponential_jitter
import hashlib
import json
import redis
r = redis.Redis()
def cache_key(model, messages, params):
payload = {"model": model, "messages": messages, "params": params}
raw = json.dumps(payload, sort_keys=True).encode()
return "llm:" + hashlib.sha256(raw).hexdigest()
@retry(
stop=stop_after_attempt(3),
wait=wait_exponential_jitter(initial=1, max=10),
)
def call_llm(client, model, messages, **params):
key = cache_key(model, messages, params)
cached = r.get(key)
if cached:
return json.loads(cached)
resp = client.chat.completions.create(
model=model,
messages=messages,
**params,
)
result = resp.model_dump()
r.setex(key, 3600, json.dumps(result))
return result
This works well if you can centralize all LLM calls in one helper.
Option 2: Put a gateway/proxy in front
If you want no app code changes or minimal changes, use an LLM gateway/proxy that supports:
- response caching
- automatic retries
- request logging
- fallback providers
- budgets/rate limiting
Examples of categories of tools:
- API gateways / LLM proxies
- observability middleware
- self-hosted request routers
Your app points to the proxy instead of the vendor endpoint. The proxy forwards requests to the real model provider.
Why this is best
- No rewriting your app’s business logic
- Centralized policies
- Easier to add multiple providers later
- Better for teams and production environments
Option 3: Use HTTP middleware if your SDK supports it
Some SDKs let you attach:
- retry handlers
- custom transport adapters
- request/response interceptors
This is ideal when you want to keep the same SDK but add cross-cutting behavior.
Caching tips for LLMs
LLM caching is trickier than normal HTTP caching because outputs can be sensitive to:
- model name/version
- temperature/top_p
- system prompt
- all messages in the conversation
- tool/function definitions
- streaming vs non-streaming
- user/session-specific context
Good cache key inputs
Include:
- model
- full message list
- decoding params
- tool schema
- system prompt
- any retrieved context
- tenant/user scope if needed
Avoid caching when:
- temperature is high and you need unique outputs
- requests depend on live data
- outputs are personalized or sensitive
- tool calls must always be fresh
A common compromise:
- cache only deterministic requests, e.g.
temperature=0 - use short TTLs
- cache only “safe” endpoints like summarization or classification
Retry tips
Retry only on transient errors:
- 408 request timeout
- 429 rate limited
- 500/502/503/504
- network timeouts / connection resets
Avoid retrying:
- validation errors
- auth failures
- bad prompts
- quota exhaustion unless you expect a brief recovery
Use:
- exponential backoff
- jitter
- max attempts
- circuit breaker if failures persist
A practical minimal setup
If you want the least invasive approach:
- Create one LLM client wrapper
- Add:
- Redis cache
- retry policy
- TTL
- Point all app code to that wrapper
- Later, move to a gateway if needed
If you want truly “without rewriting”
Use a proxy that speaks the same API as the provider:
- app keeps the same request format
- just change the base URL / endpoint
- proxy adds caching and retries
That’s the closest thing to plug-and-play.
If you want, I can show:
- a Python implementation with Redis + retries,
- a Node.js version, or
- a gateway/proxy architecture that requires only an endpoint change.
Brands