Prompt
Need LLM proxy with caching and rate limiting
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
Here’s a practical design for an LLM proxy with caching and rate limiting.
What it should do
- Accept requests from your apps/users
- Forward them to one or more LLM providers
- Cache repeatable responses
- Enforce per-user / per-key / per-org rate limits
- Log usage, latency, and cost
- Optionally retry/fallback across models
Recommended architecture
Core flow
- Client → Proxy API
- Auth / tenant identification
- Rate limit check
- Cache lookup
- If cache hit → return response
- If cache miss → call LLM provider
- Store response in cache
- Record metrics and usage
- Return response
Components
1) API Gateway / Proxy service
Could be:
- FastAPI / Node.js / Go service
- Handles request normalization and streaming support
- Routes to OpenAI, Anthropic, Azure OpenAI, local models, etc.
2) Cache
Use Redis for:
- Fast lookup
- TTL expiration
- Distributed deployment
Cache key should include:
- provider/model
- system prompt
- user prompt/messages
- temperature/top_p/max_tokens
- tools/function schema
- tenant/user identifier if needed
- any other parameters affecting output
Example cache key:
llm:{tenant}:{provider}:{model}:{hash(normalized_request)}
3) Rate limiting
Use Redis-based:
- Token bucket
- Sliding window
- Fixed window if simpler
Limit dimensions:
- per API key
- per user
- per organization
- per model/provider
- global safety limit
Example policies:
- 60 requests/minute per user
- 1,000 requests/hour per org
- 10 concurrent requests per key
4) Request normalization
Normalize requests before caching/rate limiting:
- Sort message fields
- Remove non-semantic fields
- Canonicalize whitespace if appropriate
- Exclude timestamps/request IDs from cache hash
5) Observability
Track:
- cache hit rate
- request latency
- provider error rate
- token usage
- estimated cost
- rate-limit rejects
Caching strategy
Good candidates for caching
- Deterministic or low-temperature requests
- Embedding requests
- Tool-free completions
- Summarization with fixed input
- System prompt + prompt templates
Avoid caching or be careful with
- Very high temperature outputs
- Requests containing user-specific private data
- Streaming partials unless you cache final full result
- Requests with tool calls / external side effects
- Sensitive prompts unless explicitly allowed
Cache TTL ideas
- Embeddings: hours to days
- Chat completions: minutes to hours
- Code generation: maybe shorter TTL
- Per tenant configurable TTL
Rate limiting strategy
Example policies
- Per key: 100 requests/min
- Per org: 5,000 requests/day
- Per model: stricter on expensive models
- Burst allowance: e.g. 20 requests instant burst, 100/min sustained
Concurrency control
Useful to prevent thundering herd:
- Max in-flight requests per tenant
- Queue or reject when exceeded
Minimal endpoint design
POST /v1/chat/completions
Proxy-compatible with OpenAI-style APIs.
Headers:
Authorization: Bearer <api_key>- optional
X-Tenant-ID
Request:
{
"model": "gpt-4.1-mini",
"messages": [
{"role": "system", "content": "You are helpful"},
{"role": "user", "content": "Explain caching"}
],
"temperature": 0.2
}
Response:
- Return provider-compatible JSON
- Add headers like:
X-Cache: HIT|MISSX-RateLimit-RemainingX-RateLimit-Reset
Implementation sketch
Pseudocode
def handle_request(req):
tenant = authenticate(req)
enforce_rate_limit(tenant, req)
normalized = normalize(req)
key = cache_key(tenant, normalized)
cached = redis.get(key)
if cached:
return cached, headers={"X-Cache": "HIT"}
resp = call_provider(req)
redis.setex(key, ttl_for(req), resp)
log_usage(tenant, req, resp)
return resp, headers={"X-Cache": "MISS"}
Recommended tech stack
Simple and solid
- FastAPI or Node.js/Express
- Redis for cache + rate limiting
- PostgreSQL for audit logs and usage history
- Prometheus + Grafana for metrics
- OpenTelemetry for tracing
Deployment
- Docker + Kubernetes, or
- ECS/Fargate, or
- Cloud Run
Important security considerations
- Do not cache secrets unless absolutely intended
- Encrypt logs or redact sensitive prompt content
- Separate cache namespaces per tenant
- Store API keys securely
- Prevent cache poisoning by validating request structure
- Add request signing if clients are untrusted
If you want a quick starting point
I can provide one of these:
- FastAPI implementation
- Node.js/Express implementation
- Go implementation
- Redis schema + rate limiting algorithm
- Production architecture diagram
- OpenAI-compatible proxy example
If you want, I can generate a working FastAPI proxy with Redis caching and token-bucket rate limiting next.