Prompt

How do I add caching and retries to LLM requests without rewriting my app?

Artificial Intelligence / AI Infrastructure1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

The easiest way is to put a wrapper/proxy layer between your app and the LLM provider, so your app keeps calling the same API shape while the wrapper handles caching, retries, rate limits, logging, and fallback behavior.

Option 1: Use a drop-in client wrapper

If your app already uses a standard SDK, wrap the call with a small helper:

  • Retries for transient failures: timeouts, 429s, 5xx
  • Cache for repeated prompts
  • Backoff + jitter to avoid retry storms

Example pattern

from tenacity import retry, stop_after_attempt, wait_exponential_jitter
import hashlib
import json
import redis

r = redis.Redis()

def cache_key(model, messages, params):
    payload = {"model": model, "messages": messages, "params": params}
    raw = json.dumps(payload, sort_keys=True).encode()
    return "llm:" + hashlib.sha256(raw).hexdigest()

@retry(
    stop=stop_after_attempt(3),
    wait=wait_exponential_jitter(initial=1, max=10),
)
def call_llm(client, model, messages, **params):
    key = cache_key(model, messages, params)
    cached = r.get(key)
    if cached:
        return json.loads(cached)

    resp = client.chat.completions.create(
        model=model,
        messages=messages,
        **params,
    )
    result = resp.model_dump()

    r.setex(key, 3600, json.dumps(result))
    return result

This works well if you can centralize all LLM calls in one helper.


Option 2: Put a gateway/proxy in front

If you want no app code changes or minimal changes, use an LLM gateway/proxy that supports:

  • response caching
  • automatic retries
  • request logging
  • fallback providers
  • budgets/rate limiting

Examples of categories of tools:

  • API gateways / LLM proxies
  • observability middleware
  • self-hosted request routers

Your app points to the proxy instead of the vendor endpoint. The proxy forwards requests to the real model provider.

Why this is best

  • No rewriting your app’s business logic
  • Centralized policies
  • Easier to add multiple providers later
  • Better for teams and production environments

Option 3: Use HTTP middleware if your SDK supports it

Some SDKs let you attach:

  • retry handlers
  • custom transport adapters
  • request/response interceptors

This is ideal when you want to keep the same SDK but add cross-cutting behavior.


Caching tips for LLMs

LLM caching is trickier than normal HTTP caching because outputs can be sensitive to:

  • model name/version
  • temperature/top_p
  • system prompt
  • all messages in the conversation
  • tool/function definitions
  • streaming vs non-streaming
  • user/session-specific context

Good cache key inputs

Include:

  • model
  • full message list
  • decoding params
  • tool schema
  • system prompt
  • any retrieved context
  • tenant/user scope if needed

Avoid caching when:

  • temperature is high and you need unique outputs
  • requests depend on live data
  • outputs are personalized or sensitive
  • tool calls must always be fresh

A common compromise:

  • cache only deterministic requests, e.g. temperature=0
  • use short TTLs
  • cache only “safe” endpoints like summarization or classification

Retry tips

Retry only on transient errors:

  • 408 request timeout
  • 429 rate limited
  • 500/502/503/504
  • network timeouts / connection resets

Avoid retrying:

  • validation errors
  • auth failures
  • bad prompts
  • quota exhaustion unless you expect a brief recovery

Use:

  • exponential backoff
  • jitter
  • max attempts
  • circuit breaker if failures persist

A practical minimal setup

If you want the least invasive approach:

  1. Create one LLM client wrapper
  2. Add:
    • Redis cache
    • retry policy
    • TTL
  3. Point all app code to that wrapper
  4. Later, move to a gateway if needed

If you want truly “without rewriting”

Use a proxy that speaks the same API as the provider:

  • app keeps the same request format
  • just change the base URL / endpoint
  • proxy adds caching and retries

That’s the closest thing to plug-and-play.

If you want, I can show:

  1. a Python implementation with Redis + retries,
  2. a Node.js version, or
  3. a gateway/proxy architecture that requires only an endpoint change.

Brands

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.