Prompt

How do I handle rate limits across OpenAI, Anthropic, and open-source models?

Artificial Intelligence / AI Infrastructure1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

A good way to handle rate limits across OpenAI, Anthropic, and open-source models is to treat them as one shared “request budget” problem, then add provider-specific adapters.

1) Normalize everything into one limiter

Track per request:

  • Requests per minute/hour
  • Tokens per minute/hour
  • Concurrency
  • Cost budget
  • Latency SLO

Even if each provider exposes different limits, your app should convert them into a common internal policy like:

  • global max concurrent LLM calls
  • per-provider max QPS
  • per-model token budget
  • fallback priority order

2) Use hierarchical throttling

Set limits at 3 levels:

Global

Protect your app from overload:

  • total outbound LLM concurrency
  • total token throughput
  • total retry budget

Provider

Each provider gets its own bucket:

  • OpenAI: separate request/token buckets
  • Anthropic: separate request/token buckets
  • Open-source inference: usually GPU queue depth / concurrent jobs

Model

Some models are more constrained than others:

  • expensive models get lower priority
  • premium users get reserved capacity
  • long-context requests get their own budget

3) Use exponential backoff with jitter

When you hit a 429 or throttling error:

  • retry with exponential backoff
  • add random jitter
  • respect provider retry headers if present

Example policy:

  • 1st retry: 250–500 ms
  • 2nd: 500–1000 ms
  • 3rd: 1–2 s
  • cap at 10–30 s

Avoid synchronized retry storms.

4) Distinguish hard vs soft failures

Not all rate limit issues should be retried the same way.

Retryable

  • HTTP 429
  • transient 5xx
  • queue saturation on self-hosted models

Non-retryable or needs different action

  • quota exhausted
  • invalid API key
  • context too long
  • model unavailable for that tenant

5) Add fallback routing

If one provider is saturated:

  1. try primary model
  2. fall back to cheaper/slower model
  3. fall back to open-source model
  4. degrade response quality gracefully

Examples:

  • summarization model instead of flagship model
  • smaller context window
  • cached response
  • partial answer with “continue if needed”

6) Queue and batch when possible

For non-interactive work:

  • queue requests instead of sending immediately
  • batch embeddings or classification jobs
  • coalesce duplicate prompts
  • deduplicate identical requests using a cache key

This is especially useful for open-source models where batching can improve GPU utilization.

7) Use token-aware scheduling

Token usage, not just request count, usually causes rate issues.

Prioritize:

  • short prompts first
  • small completions before long generations
  • streaming responses if they reduce tail latency
  • truncation/summarization for oversized inputs

A simple strategy:

  • estimate prompt + completion tokens before dispatch
  • reject, defer, or downshift if request exceeds remaining budget

8) Implement a circuit breaker

If a provider starts failing repeatedly:

  • open the circuit after N failures
  • stop sending traffic briefly
  • route to backup provider
  • periodically probe for recovery

This prevents cascading failures and wasted retries.

9) Separate user-facing and background traffic

Reserve capacity for:

  • interactive chat
  • critical API calls

Throttle or defer:

  • offline jobs
  • analytics
  • bulk imports
  • evaluation runs

This prevents background workloads from exhausting limits.

10) Monitor the right metrics

Track by provider and model:

  • 429 rate
  • retry count
  • token throughput
  • queue wait time
  • success latency
  • fallback rate
  • cost per request
  • GPU utilization for open-source models

Alert when:

  • retry rate spikes
  • queue backlog grows
  • latency hits SLO
  • a provider’s effective throughput drops

11) Open-source models need their own “rate limit”

Even though they don’t have API quotas, they still have capacity limits:

  • GPU memory
  • max concurrent generations
  • batch size
  • context length
  • worker queue depth

Use:

  • a job queue
  • admission control
  • autoscaling if possible
  • per-model concurrency caps

12) Practical architecture

A solid setup looks like:

  • LLM Gateway

    • chooses provider/model
    • applies rate limits
    • tracks budgets
    • handles retries/fallbacks
  • Policy engine

    • per-tenant limits
    • priority rules
    • cost thresholds
  • Provider adapters

    • OpenAI adapter
    • Anthropic adapter
    • open-source adapter
  • Shared observability

    • metrics
    • logs
    • traces

13) Simple rule set you can start with

  • 5–10% of capacity reserved for retries
  • exponential backoff with jitter
  • max 2–3 retries for interactive requests
  • no infinite retries
  • global concurrency cap
  • per-provider token budgets
  • fallback after first hard 429 or after 2 transient failures
  • circuit break after repeated provider errors

14) Pseudocode sketch

def call_llm(request):
    provider = choose_provider(request)

    if not budget_available(provider, request):
        provider = fallback_provider(request)

    try:
        return provider.send(request)
    except RateLimitError:
        record_rate_limit(provider)
        sleep(backoff_with_jitter())
        if can_retry(request):
            return call_llm(request)
        return fallback_or_fail(request)
    except ProviderOverloaded:
        return route_to_queue_or_fallback(request)

15) Key design principle

Don’t treat OpenAI, Anthropic, and open-source as separate one-off integrations.
Treat them as capacity-constrained backends behind a single policy layer.

If you want, I can also give you:

  1. a concrete Python rate-limit manager,
  2. a Redis-based distributed limiter design, or
  3. a provider comparison table for OpenAI vs Anthropic vs self-hosted models.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.