Prompt

How do I cache LLM responses to reduce cost and latency?

Artificial Intelligence / AI Infrastructure1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

Caching LLM responses can save a lot of money and cut latency, but it works best when you cache the right layer and have a good invalidation strategy.

What to cache

1) Exact prompt/response cache

Store the model output keyed by:

  • model name/version
  • full prompt/messages
  • decoding params (temperature, top_p, max_tokens, etc.)
  • tool/function schema, if used
  • system prompt version

Use this when you expect repeated identical requests.

Pros: simplest, biggest savings for repeated traffic
Cons: low hit rate if prompts vary even slightly


2) Semantic cache

Cache responses by meaning, not exact text. For a new query:

  1. embed the prompt
  2. retrieve similar past prompts
  3. if similarity is high enough, reuse or adapt the prior answer

Best for customer support, FAQ, internal knowledge assistants.

Pros: higher hit rate
Cons: harder to implement; risk of incorrect reuse


3) Prefix / prompt fragment cache

If your prompts have a stable prefix (system instructions, policy text, long context), cache that prefix or use provider-supported prompt caching when available.

Best when the first part of the prompt is long and reused across many calls.

Pros: reduces repeated token processing
Cons: depends on provider/model support


Recommended architecture

Fast path

  1. Build a deterministic cache key from:
    • normalized messages
    • model
    • parameters
    • tool definitions
  2. Check Redis/memory cache
  3. If hit, return cached response
  4. If miss, call LLM
  5. Store result with TTL

Optional semantic fallback

If exact cache miss:

  1. search vector DB for similar prompts
  2. if above threshold, return cached answer or use it as context
  3. else call LLM

Cache key design

A good cache key should include everything that can affect output:

  • model: gpt-4.1-mini
  • messages: canonical JSON serialization
  • temperature, top_p, max_tokens
  • system_prompt_version
  • tools / function schemas
  • retrieval corpus version if using RAG
  • tenant/user id if responses are personalized

Example key:

llm:v3:model=gpt-4.1-mini:temp=0:hash=SHA256(canonical_request)

Important: normalize inputs first:

  • trim whitespace
  • stable JSON ordering
  • remove irrelevant metadata
  • canonicalize tool schemas

Storage choices

In-memory cache

  • fastest
  • good for single instance or local hot cache
  • not durable

Redis / KeyDB

  • common production choice
  • shared across workers
  • supports TTL, LRU-like eviction

Database

  • good for audit/history
  • slower than Redis
  • usually not ideal as the primary cache

Vector DB

  • for semantic caching
  • store embeddings + response metadata

Invalidation strategy

LLM caches go stale when:

  • system prompt changes
  • tools change
  • underlying knowledge base changes
  • user-specific state changes

Use:

  • TTL for freshness
  • versioned keys so updates invalidate automatically
  • tenant scoping for personalized data
  • cache bypass for critical or sensitive queries

Practical policies

Cache only when deterministic

Use caching for:

  • temperature=0 or low temperature
  • classification
  • extraction
  • FAQs
  • summarization of stable inputs

Avoid strict reuse for:

  • creative writing
  • open-ended brainstorming
  • anything requiring freshness

Cache the expensive part

If you use RAG:

  • cache retrieved documents separately
  • cache final answer only if retrieval set is stable
  • consider caching embeddings for repeated prompts

Set TTL based on content type

  • FAQ: hours to days
  • internal docs: minutes to hours
  • user-specific answers: short TTL or no cache
  • policy/compliance outputs: very conservative

Example implementation pattern

import hashlib
import json
import redis

r = redis.Redis()

def canonicalize_request(req):
    return json.dumps(req, sort_keys=True, separators=(",", ":"))

def cache_key(req):
    s = canonicalize_request(req)
    h = hashlib.sha256(s.encode()).hexdigest()
    return f"llm:v1:{h}"

def get_llm_response(req, call_llm):
    key = cache_key(req)
    cached = r.get(key)
    if cached:
        return json.loads(cached)

    resp = call_llm(req)
    r.setex(key, 3600, json.dumps(resp))  # 1 hour TTL
    return resp

Cost/latency best practices

  • Cache at the application layer before calling the model
  • Use streaming only after cache check
  • Deduplicate concurrent identical requests with request coalescing
  • Cache embeddings if using semantic search
  • Monitor hit rate, TTL, and stale-answer rate

Things to watch out for

  • Personal data leakage: never reuse across users unless safe
  • Prompt injection: don’t cache unsafe tool outputs blindly
  • Stale answers: especially with RAG or rapidly changing info
  • Over-caching: low hit rate can add complexity without benefit

Quick recommendation

If you want the best ROI:

  1. Start with exact response caching in Redis
  2. Use temperature=0 for cacheable endpoints
  3. Add versioned keys + TTL
  4. If needed, add semantic caching for FAQ/support use cases

If you want, I can show you:

  • a Redis-based caching design,
  • a semantic cache example,
  • or how to do this for a specific SDK/framework.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.