Prompt
I'm building a product that needs model fallback and rate-limit handling; what gateway stack makes sense?
Latest observation
Oct 1, 2026
OpenAI APIWeb search: off
If you need model fallback + rate-limit handling, the cleanest approach is usually:
Recommended gateway stack
1) LiteLLM Proxy as the LLM gateway
Best fit if you want one API in front of multiple providers/models.
Why it makes sense:
- Model fallback / routing across OpenAI, Anthropic, Azure, Gemini, local models, etc.
- Retries and failover when a provider/model is down or rate-limited
- Rate limiting / quotas per key, team, tenant, or route
- Cost tracking and usage logging
- OpenAI-compatible API, which makes app integration easier
Typical pattern:
- Your app talks to LiteLLM Proxy
- Proxy routes to primary model
- On 429/5xx/timeouts, it falls back to secondary models/providers
- It can also enforce per-user / per-org limits
2) Put an API gateway in front for auth and traffic shaping
If you need broader API management, pair LiteLLM with:
- Kong
- Envoy
- NGINX
- Cloudflare API Gateway
- AWS API Gateway
Use this layer for:
- AuthN/AuthZ
- IP allowlists
- Global rate limiting
- WAF / bot protection
- Request size limits
- Tenant isolation
A common stack is:
- Client → API Gateway (Kong/Cloudflare/NGINX) → LiteLLM Proxy → model providers
3) Add a queue/workflow layer for non-interactive jobs
If some requests can be async, use:
- Celery
- RQ
- BullMQ
- SQS + workers
- Temporal
This helps when:
- Providers are rate-limited
- You want automatic backoff
- You need durable retries for long-running tasks
What I’d choose by use case
Fastest path to production
- LiteLLM Proxy only
- Add provider-specific fallback rules
- Use built-in rate limits/quotas
More enterprise / multi-tenant
- Cloudflare or Kong in front
- LiteLLM Proxy behind it
- Central auth, per-tenant policies, and observability
Heavy async workloads
- API Gateway + LiteLLM + queue/worker system
- Better for batch generation, document processing, and background jobs
Important implementation details
Look for support for:
- Fallback chains by model, provider, region, or cost tier
- Exponential backoff + jitter
- Circuit breakers for bad providers
- Per-tenant rate limits
- Usage accounting
- Streaming response passthrough
- Observability: logs, traces, metrics
- Secret management for provider keys
Practical recommendation
If your main need is LLM routing with fallback and rate limits, start with:
LiteLLM Proxy + Redis (for shared limits/state) + your existing API gateway if you already have one.
If you want, I can also give you:
- a reference architecture diagram,
- a Kong vs Cloudflare vs NGINX vs LiteLLM comparison, or
- a sample fallback configuration for OpenAI/Anthropic.