Prompt
How do I cache LLM responses to reduce cost and latency?
Latest observation
Caching LLM responses can save a lot of money and cut latency, but it works best when you cache the right layer and have a good invalidation strategy.
What to cache
1) Exact prompt/response cache
Store the model output keyed by:
- model name/version
- full prompt/messages
- decoding params (
temperature,top_p,max_tokens, etc.) - tool/function schema, if used
- system prompt version
Use this when you expect repeated identical requests.
Pros: simplest, biggest savings for repeated traffic
Cons: low hit rate if prompts vary even slightly
2) Semantic cache
Cache responses by meaning, not exact text. For a new query:
- embed the prompt
- retrieve similar past prompts
- if similarity is high enough, reuse or adapt the prior answer
Best for customer support, FAQ, internal knowledge assistants.
Pros: higher hit rate
Cons: harder to implement; risk of incorrect reuse
3) Prefix / prompt fragment cache
If your prompts have a stable prefix (system instructions, policy text, long context), cache that prefix or use provider-supported prompt caching when available.
Best when the first part of the prompt is long and reused across many calls.
Pros: reduces repeated token processing
Cons: depends on provider/model support
Recommended architecture
Fast path
- Build a deterministic cache key from:
- normalized messages
- model
- parameters
- tool definitions
- Check Redis/memory cache
- If hit, return cached response
- If miss, call LLM
- Store result with TTL
Optional semantic fallback
If exact cache miss:
- search vector DB for similar prompts
- if above threshold, return cached answer or use it as context
- else call LLM
Cache key design
A good cache key should include everything that can affect output:
model:gpt-4.1-minimessages: canonical JSON serializationtemperature,top_p,max_tokenssystem_prompt_versiontools/ function schemasretrieval corpus versionif using RAGtenant/user idif responses are personalized
Example key:
llm:v3:model=gpt-4.1-mini:temp=0:hash=SHA256(canonical_request)
Important: normalize inputs first:
- trim whitespace
- stable JSON ordering
- remove irrelevant metadata
- canonicalize tool schemas
Storage choices
In-memory cache
- fastest
- good for single instance or local hot cache
- not durable
Redis / KeyDB
- common production choice
- shared across workers
- supports TTL, LRU-like eviction
Database
- good for audit/history
- slower than Redis
- usually not ideal as the primary cache
Vector DB
- for semantic caching
- store embeddings + response metadata
Invalidation strategy
LLM caches go stale when:
- system prompt changes
- tools change
- underlying knowledge base changes
- user-specific state changes
Use:
- TTL for freshness
- versioned keys so updates invalidate automatically
- tenant scoping for personalized data
- cache bypass for critical or sensitive queries
Practical policies
Cache only when deterministic
Use caching for:
temperature=0or low temperature- classification
- extraction
- FAQs
- summarization of stable inputs
Avoid strict reuse for:
- creative writing
- open-ended brainstorming
- anything requiring freshness
Cache the expensive part
If you use RAG:
- cache retrieved documents separately
- cache final answer only if retrieval set is stable
- consider caching embeddings for repeated prompts
Set TTL based on content type
- FAQ: hours to days
- internal docs: minutes to hours
- user-specific answers: short TTL or no cache
- policy/compliance outputs: very conservative
Example implementation pattern
import hashlib
import json
import redis
r = redis.Redis()
def canonicalize_request(req):
return json.dumps(req, sort_keys=True, separators=(",", ":"))
def cache_key(req):
s = canonicalize_request(req)
h = hashlib.sha256(s.encode()).hexdigest()
return f"llm:v1:{h}"
def get_llm_response(req, call_llm):
key = cache_key(req)
cached = r.get(key)
if cached:
return json.loads(cached)
resp = call_llm(req)
r.setex(key, 3600, json.dumps(resp)) # 1 hour TTL
return resp
Cost/latency best practices
- Cache at the application layer before calling the model
- Use streaming only after cache check
- Deduplicate concurrent identical requests with request coalescing
- Cache embeddings if using semantic search
- Monitor hit rate, TTL, and stale-answer rate
Things to watch out for
- Personal data leakage: never reuse across users unless safe
- Prompt injection: don’t cache unsafe tool outputs blindly
- Stale answers: especially with RAG or rapidly changing info
- Over-caching: low hit rate can add complexity without benefit
Quick recommendation
If you want the best ROI:
- Start with exact response caching in Redis
- Use temperature=0 for cacheable endpoints
- Add versioned keys + TTL
- If needed, add semantic caching for FAQ/support use cases
If you want, I can show you:
- a Redis-based caching design,
- a semantic cache example,
- or how to do this for a specific SDK/framework.