Prompt
I'm unhappy with our current setup for retries and rate limits on LLM calls; what should I use?
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
For LLM calls, I’d recommend using a single wrapper that handles both retries and rate limiting, rather than sprinkling retry logic and sleep calls throughout your code.
Good default choice
If you’re in Python, a strong setup is:
tenacityfor retriesasynciosemaphore / token bucket / leaky bucket for concurrency and rate limiting- optionally a small shared helper module so every LLM call goes through the same policy
If you want an off-the-shelf approach, use a framework or SDK layer that already supports:
- exponential backoff
- jitter
- max retry count
- retry only on transient errors
- request pacing / concurrency limits
- circuit breaking or fallback behavior
What to retry
Retry only transient failures, like:
- HTTP 429 rate limit
- 500/502/503/504
- timeouts
- network disconnects
Do not blindly retry:
- invalid requests
- auth failures
- bad prompts / schema errors
- context-length errors
- deterministic application errors
What to use for rate limiting
Prefer one of these:
- Concurrency limit: cap how many requests are in flight at once
- Token bucket: cap request rate over time
- Both together: best in practice for LLM APIs
Why both?
A concurrency cap prevents stampedes; a token bucket keeps you under provider RPM/TPM limits.
Recommended pattern
A solid production pattern is:
- acquire rate-limit permission
- call the model
- if transient error, retry with exponential backoff + jitter
- honor
Retry-Afterheaders if the provider returns them - log attempt count, latency, and final outcome
If you want a higher-level answer
If you’re asking “what should I use” in a team setting, my practical recommendation is:
- Python:
tenacity+aiolimiteror an async semaphore - Node: a retry library like
p-retryplus a rate limiter likebottleneck - Distributed systems: put limits in a shared service or gateway, not only in each app instance
My opinionated recommendation
For most teams:
- start with provider-native retries if they’re solid
- add your own retry wrapper for consistent behavior
- add centralized rate limiting if multiple workers/services share quota
If you want, I can suggest a concrete setup for your stack:
- Python sync
- Python async
- Node.js
- distributed microservices
- OpenAI / Anthropic / Azure OpenAI specific