Prompt

Why are my LLM requests getting rate limited even though I have a gateway?

Artificial Intelligence / AI Infrastructure1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

A gateway can help, but it doesn’t automatically eliminate rate limits. Common reasons you’re still getting limited:

  1. The gateway has its own limits

    • Your gateway may enforce per-user, per-org, per-key, or per-minute quotas.
    • Even if the upstream provider allows more, the gateway can still throttle you.
  2. The upstream model provider is still the bottleneck

    • Some gateways proxy requests but don’t increase your actual provider quota.
    • If the gateway forwards to a single provider key/org, you can still hit the provider’s RPM/TPM limits.
  3. You’re exceeding token-based limits, not just request counts

    • Many systems rate limit by:
      • requests per minute
      • tokens per minute
      • concurrent requests
    • A few large prompts or long outputs can trip token limits quickly.
  4. Burst traffic

    • You may be under the average limit but sending traffic in spikes.
    • Gateways often smooth bursts a bit, but not always enough.
  5. Concurrency limits

    • Too many simultaneous requests can trigger throttling even if total volume is fine.
  6. Multiple apps/clients share the same quota

    • If several services, environments, or users use the same gateway/API key, they all consume the same pool.
  7. Retry storms

    • If your client auto-retries aggressively, failed requests can multiply traffic and make rate limiting worse.
  8. Gateway misconfiguration

    • Possible issues include:
      • wrong routing to a single upstream key
      • no backoff
      • mis-set org/project limits
      • incorrect tenant identification
      • caching disabled when expected
  9. Provider-side adaptive throttling

    • Some providers apply dynamic limits based on load, account trust, abuse detection, or model tier.
  10. You’re hitting model-specific limits

  • Different models often have different quotas. A gateway won’t change that unless it routes to a model or provider with more capacity.

What to check

  • Gateway logs: is the 429 coming from the gateway or upstream?
  • Whether the limit is RPM, TPM, concurrency, or burst
  • Which API key/org/project is actually being used
  • Whether retries are amplifying traffic
  • Per-model quotas and routing rules
  • Any shared usage across environments or tenants

Practical fixes

  • Add exponential backoff with jitter
  • Reduce max concurrency
  • Batch or cache requests where possible
  • Lower prompt/output token sizes
  • Spread load across multiple keys/projects/models if allowed
  • Increase quotas with the gateway/provider
  • Make sure the gateway is actually load balancing across upstreams if intended

If you want, I can help you diagnose it from:

  • the exact 429/error message,
  • your gateway setup,
  • and whether you’re using OpenAI, Anthropic, Azure, Bedrock, etc.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.