Prompt
How do I handle rate limits across OpenAI, Anthropic, and open-source models?
Latest observation
A good way to handle rate limits across OpenAI, Anthropic, and open-source models is to treat them as one shared “request budget” problem, then add provider-specific adapters.
1) Normalize everything into one limiter
Track per request:
- Requests per minute/hour
- Tokens per minute/hour
- Concurrency
- Cost budget
- Latency SLO
Even if each provider exposes different limits, your app should convert them into a common internal policy like:
- global max concurrent LLM calls
- per-provider max QPS
- per-model token budget
- fallback priority order
2) Use hierarchical throttling
Set limits at 3 levels:
Global
Protect your app from overload:
- total outbound LLM concurrency
- total token throughput
- total retry budget
Provider
Each provider gets its own bucket:
- OpenAI: separate request/token buckets
- Anthropic: separate request/token buckets
- Open-source inference: usually GPU queue depth / concurrent jobs
Model
Some models are more constrained than others:
- expensive models get lower priority
- premium users get reserved capacity
- long-context requests get their own budget
3) Use exponential backoff with jitter
When you hit a 429 or throttling error:
- retry with exponential backoff
- add random jitter
- respect provider retry headers if present
Example policy:
- 1st retry: 250–500 ms
- 2nd: 500–1000 ms
- 3rd: 1–2 s
- cap at 10–30 s
Avoid synchronized retry storms.
4) Distinguish hard vs soft failures
Not all rate limit issues should be retried the same way.
Retryable
- HTTP 429
- transient 5xx
- queue saturation on self-hosted models
Non-retryable or needs different action
- quota exhausted
- invalid API key
- context too long
- model unavailable for that tenant
5) Add fallback routing
If one provider is saturated:
- try primary model
- fall back to cheaper/slower model
- fall back to open-source model
- degrade response quality gracefully
Examples:
- summarization model instead of flagship model
- smaller context window
- cached response
- partial answer with “continue if needed”
6) Queue and batch when possible
For non-interactive work:
- queue requests instead of sending immediately
- batch embeddings or classification jobs
- coalesce duplicate prompts
- deduplicate identical requests using a cache key
This is especially useful for open-source models where batching can improve GPU utilization.
7) Use token-aware scheduling
Token usage, not just request count, usually causes rate issues.
Prioritize:
- short prompts first
- small completions before long generations
- streaming responses if they reduce tail latency
- truncation/summarization for oversized inputs
A simple strategy:
- estimate prompt + completion tokens before dispatch
- reject, defer, or downshift if request exceeds remaining budget
8) Implement a circuit breaker
If a provider starts failing repeatedly:
- open the circuit after N failures
- stop sending traffic briefly
- route to backup provider
- periodically probe for recovery
This prevents cascading failures and wasted retries.
9) Separate user-facing and background traffic
Reserve capacity for:
- interactive chat
- critical API calls
Throttle or defer:
- offline jobs
- analytics
- bulk imports
- evaluation runs
This prevents background workloads from exhausting limits.
10) Monitor the right metrics
Track by provider and model:
- 429 rate
- retry count
- token throughput
- queue wait time
- success latency
- fallback rate
- cost per request
- GPU utilization for open-source models
Alert when:
- retry rate spikes
- queue backlog grows
- latency hits SLO
- a provider’s effective throughput drops
11) Open-source models need their own “rate limit”
Even though they don’t have API quotas, they still have capacity limits:
- GPU memory
- max concurrent generations
- batch size
- context length
- worker queue depth
Use:
- a job queue
- admission control
- autoscaling if possible
- per-model concurrency caps
12) Practical architecture
A solid setup looks like:
-
LLM Gateway
- chooses provider/model
- applies rate limits
- tracks budgets
- handles retries/fallbacks
-
Policy engine
- per-tenant limits
- priority rules
- cost thresholds
-
Provider adapters
- OpenAI adapter
- Anthropic adapter
- open-source adapter
-
Shared observability
- metrics
- logs
- traces
13) Simple rule set you can start with
- 5–10% of capacity reserved for retries
- exponential backoff with jitter
- max 2–3 retries for interactive requests
- no infinite retries
- global concurrency cap
- per-provider token budgets
- fallback after first hard 429 or after 2 transient failures
- circuit break after repeated provider errors
14) Pseudocode sketch
def call_llm(request):
provider = choose_provider(request)
if not budget_available(provider, request):
provider = fallback_provider(request)
try:
return provider.send(request)
except RateLimitError:
record_rate_limit(provider)
sleep(backoff_with_jitter())
if can_retry(request):
return call_llm(request)
return fallback_or_fail(request)
except ProviderOverloaded:
return route_to_queue_or_fallback(request)
15) Key design principle
Don’t treat OpenAI, Anthropic, and open-source as separate one-off integrations.
Treat them as capacity-constrained backends behind a single policy layer.
If you want, I can also give you:
- a concrete Python rate-limit manager,
- a Redis-based distributed limiter design, or
- a provider comparison table for OpenAI vs Anthropic vs self-hosted models.