Prompt

Need LLM proxy with caching and rate limiting

Artificial Intelligence / AI Infrastructure1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

Here’s a practical design for an LLM proxy with caching and rate limiting.

What it should do

  • Accept requests from your apps/users
  • Forward them to one or more LLM providers
  • Cache repeatable responses
  • Enforce per-user / per-key / per-org rate limits
  • Log usage, latency, and cost
  • Optionally retry/fallback across models

Recommended architecture

Core flow

  1. Client → Proxy API
  2. Auth / tenant identification
  3. Rate limit check
  4. Cache lookup
  5. If cache hit → return response
  6. If cache miss → call LLM provider
  7. Store response in cache
  8. Record metrics and usage
  9. Return response

Components

1) API Gateway / Proxy service

Could be:

  • FastAPI / Node.js / Go service
  • Handles request normalization and streaming support
  • Routes to OpenAI, Anthropic, Azure OpenAI, local models, etc.

2) Cache

Use Redis for:

  • Fast lookup
  • TTL expiration
  • Distributed deployment

Cache key should include:

  • provider/model
  • system prompt
  • user prompt/messages
  • temperature/top_p/max_tokens
  • tools/function schema
  • tenant/user identifier if needed
  • any other parameters affecting output

Example cache key:

llm:{tenant}:{provider}:{model}:{hash(normalized_request)}

3) Rate limiting

Use Redis-based:

  • Token bucket
  • Sliding window
  • Fixed window if simpler

Limit dimensions:

  • per API key
  • per user
  • per organization
  • per model/provider
  • global safety limit

Example policies:

  • 60 requests/minute per user
  • 1,000 requests/hour per org
  • 10 concurrent requests per key

4) Request normalization

Normalize requests before caching/rate limiting:

  • Sort message fields
  • Remove non-semantic fields
  • Canonicalize whitespace if appropriate
  • Exclude timestamps/request IDs from cache hash

5) Observability

Track:

  • cache hit rate
  • request latency
  • provider error rate
  • token usage
  • estimated cost
  • rate-limit rejects

Caching strategy

Good candidates for caching

  • Deterministic or low-temperature requests
  • Embedding requests
  • Tool-free completions
  • Summarization with fixed input
  • System prompt + prompt templates

Avoid caching or be careful with

  • Very high temperature outputs
  • Requests containing user-specific private data
  • Streaming partials unless you cache final full result
  • Requests with tool calls / external side effects
  • Sensitive prompts unless explicitly allowed

Cache TTL ideas

  • Embeddings: hours to days
  • Chat completions: minutes to hours
  • Code generation: maybe shorter TTL
  • Per tenant configurable TTL

Rate limiting strategy

Example policies

  • Per key: 100 requests/min
  • Per org: 5,000 requests/day
  • Per model: stricter on expensive models
  • Burst allowance: e.g. 20 requests instant burst, 100/min sustained

Concurrency control

Useful to prevent thundering herd:

  • Max in-flight requests per tenant
  • Queue or reject when exceeded

Minimal endpoint design

POST /v1/chat/completions

Proxy-compatible with OpenAI-style APIs.

Headers:

  • Authorization: Bearer <api_key>
  • optional X-Tenant-ID

Request:

{
  "model": "gpt-4.1-mini",
  "messages": [
    {"role": "system", "content": "You are helpful"},
    {"role": "user", "content": "Explain caching"}
  ],
  "temperature": 0.2
}

Response:

  • Return provider-compatible JSON
  • Add headers like:
    • X-Cache: HIT|MISS
    • X-RateLimit-Remaining
    • X-RateLimit-Reset

Implementation sketch

Pseudocode

def handle_request(req):
    tenant = authenticate(req)
    enforce_rate_limit(tenant, req)

    normalized = normalize(req)
    key = cache_key(tenant, normalized)

    cached = redis.get(key)
    if cached:
        return cached, headers={"X-Cache": "HIT"}

    resp = call_provider(req)
    redis.setex(key, ttl_for(req), resp)

    log_usage(tenant, req, resp)
    return resp, headers={"X-Cache": "MISS"}

Recommended tech stack

Simple and solid

  • FastAPI or Node.js/Express
  • Redis for cache + rate limiting
  • PostgreSQL for audit logs and usage history
  • Prometheus + Grafana for metrics
  • OpenTelemetry for tracing

Deployment

  • Docker + Kubernetes, or
  • ECS/Fargate, or
  • Cloud Run

Important security considerations

  • Do not cache secrets unless absolutely intended
  • Encrypt logs or redact sensitive prompt content
  • Separate cache namespaces per tenant
  • Store API keys securely
  • Prevent cache poisoning by validating request structure
  • Add request signing if clients are untrusted

If you want a quick starting point

I can provide one of these:

  1. FastAPI implementation
  2. Node.js/Express implementation
  3. Go implementation
  4. Redis schema + rate limiting algorithm
  5. Production architecture diagram
  6. OpenAI-compatible proxy example

If you want, I can generate a working FastAPI proxy with Redis caching and token-bucket rate limiting next.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.