Prompt
Why are my LLM requests getting rate limited even though I have a gateway?
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
A gateway can help, but it doesn’t automatically eliminate rate limits. Common reasons you’re still getting limited:
-
The gateway has its own limits
- Your gateway may enforce per-user, per-org, per-key, or per-minute quotas.
- Even if the upstream provider allows more, the gateway can still throttle you.
-
The upstream model provider is still the bottleneck
- Some gateways proxy requests but don’t increase your actual provider quota.
- If the gateway forwards to a single provider key/org, you can still hit the provider’s RPM/TPM limits.
-
You’re exceeding token-based limits, not just request counts
- Many systems rate limit by:
- requests per minute
- tokens per minute
- concurrent requests
- A few large prompts or long outputs can trip token limits quickly.
- Many systems rate limit by:
-
Burst traffic
- You may be under the average limit but sending traffic in spikes.
- Gateways often smooth bursts a bit, but not always enough.
-
Concurrency limits
- Too many simultaneous requests can trigger throttling even if total volume is fine.
-
Multiple apps/clients share the same quota
- If several services, environments, or users use the same gateway/API key, they all consume the same pool.
-
Retry storms
- If your client auto-retries aggressively, failed requests can multiply traffic and make rate limiting worse.
-
Gateway misconfiguration
- Possible issues include:
- wrong routing to a single upstream key
- no backoff
- mis-set org/project limits
- incorrect tenant identification
- caching disabled when expected
- Possible issues include:
-
Provider-side adaptive throttling
- Some providers apply dynamic limits based on load, account trust, abuse detection, or model tier.
-
You’re hitting model-specific limits
- Different models often have different quotas. A gateway won’t change that unless it routes to a model or provider with more capacity.
What to check
- Gateway logs: is the 429 coming from the gateway or upstream?
- Whether the limit is RPM, TPM, concurrency, or burst
- Which API key/org/project is actually being used
- Whether retries are amplifying traffic
- Per-model quotas and routing rules
- Any shared usage across environments or tenants
Practical fixes
- Add exponential backoff with jitter
- Reduce max concurrency
- Batch or cache requests where possible
- Lower prompt/output token sizes
- Spread load across multiple keys/projects/models if allowed
- Increase quotas with the gateway/provider
- Make sure the gateway is actually load balancing across upstreams if intended
If you want, I can help you diagnose it from:
- the exact 429/error message,
- your gateway setup,
- and whether you’re using OpenAI, Anthropic, Azure, Bedrock, etc.