Prompt
What should I use for caching and rate limits on AI traffic?
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
For AI traffic, a good rule is:
- Cache: use a semantic cache for repeated or near-duplicate prompts/responses
- Rate limit: use a distributed token-bucket or sliding-window limiter backed by Redis or your API gateway
- For LLM-specific traffic: measure and limit by tokens, not just requests
Caching options
1) Prompt/response caching
Best when users ask the same or very similar questions.
Use when:
- prompts repeat often
- responses are deterministic-ish
- you can tolerate serving an earlier answer
Tools:
- Redis for fast key-value caching
- LLM semantic cache libraries/providers for embedding-based similarity matching
Good patterns:
- exact match cache:
hash(model + system_prompt + user_prompt + params) - semantic cache: compare embeddings and reuse if similarity is high
2) Response caching at the gateway
If you have a stable API and identical requests happen often, cache at:
- CDN
- API gateway
- Redis
Useful for:
- system prompts
- tool results
- retrieval results
- expensive pre-processing
3) Retrieval / embedding cache
If your app uses RAG, cache:
- embeddings
- document chunks
- retrieval query results
This can save a lot more money than caching final completions alone.
Rate limiting options
Best practice: token-aware rate limits
For AI APIs, limit:
- requests per minute
- tokens per minute
- concurrency per user/org/API key
Why:
- one request may be cheap
- another may be extremely expensive
Common algorithms
- Token bucket: best general choice
- Leaky bucket: good for smoothing bursts
- Sliding window: good for fairness and accuracy
Backends
- Redis: most common for distributed rate limiting
- API gateway / edge proxy: Cloudflare, Kong, Envoy, NGINX, Apigee
- Managed services: if you want less ops
Practical recommendation
If you want a simple, solid setup:
- Cache: Redis + optional semantic cache layer
- Rate limit: Redis-backed token bucket
- Enforce limits on:
- requests/minute
- tokens/minute
- concurrent generations
- Set separate limits for:
- per user
- per API key
- per organization
- per model tier
If you want an architecture that works well
- Edge/API gateway: auth + basic request rate limiting
- Application layer: token-aware rate limiting and caching
- Redis: shared state for both
- LLM middleware: prompt normalization, semantic cache, token accounting
Quick picks by need
- Fastest to implement: Redis + token bucket
- Best for repeated prompts: semantic cache
- Best for controlling cost: token-based limiting
- Best for production scale: gateway + Redis + token-aware policies
If you want, I can recommend a specific stack for:
- Python/FastAPI,
- Node.js, or
- Cloudflare/Kong/Envoy.