Prompt
How do I enforce rate limits and quotas on public APIs?
Latest observation
To enforce rate limits and quotas on public APIs, combine policy design, enforcement at the edge, accurate counters, and good client feedback.
1) Define the policies
Decide what you want to limit:
- Rate limit: requests per second/minute/hour
- Quota: total usage over a longer period, e.g. 100,000 requests/month
- Burst vs steady-state: allow short spikes but cap sustained traffic
- Per dimension:
- per API key
- per user/account/tenant
- per IP address
- per endpoint or method
- per plan/tier
- per geography or device, if needed
A common policy model:
- Free: 100 RPM, 10k/month
- Pro: 1,000 RPM, 1M/month
- Enterprise: custom
2) Enforce as early as possible
Best practice is to reject traffic at the API gateway / edge / load balancer before it reaches your app.
Typical layers:
- CDN / WAF
- API Gateway
- Reverse proxy
- Application middleware
- Backend service guards
Prefer the edge because it:
- reduces backend load
- is easier to centralize
- makes abuse cheaper to stop
3) Use the right algorithm
Common algorithms:
Token bucket
- Tokens accumulate at a fixed rate.
- Each request consumes a token.
- Good for allowing bursts while keeping average rate bounded.
- Very common for API rate limiting.
Leaky bucket
- Smooths traffic to a constant drain rate.
- Good for strict pacing, less bursty.
Fixed window
- Simple: count requests in a time window.
- Cheap, but can allow boundary bursts.
Sliding window
- More accurate than fixed window.
- Better fairness, more computation.
For most public APIs, token bucket is a strong default.
4) Store counters centrally
If you run multiple API servers, in-memory counters alone won’t work reliably.
Use a shared store such as:
- Redis
- Memcached
- DynamoDB / Cassandra / PostgreSQL depending on scale and consistency needs
For rate limits, Redis is often used because it’s fast and supports atomic operations.
Important:
- Counters must be atomic
- Updates should be race-safe
- Use TTLs to expire window keys automatically
5) Make the limit decision atomically
A request should:
- Identify caller: API key / user / tenant / IP
- Look up policy
- Atomically check and increment counters
- Allow or reject
- Return remaining quota info
If the check and increment aren’t atomic, concurrent requests can exceed the limit.
6) Return clear headers and status codes
When a client is rate-limited, respond with:
- HTTP 429 Too Many Requests
Include helpful headers such as:
Retry-AfterX-RateLimit-LimitX-RateLimit-RemainingX-RateLimit-Reset
If you use standard or near-standard headers, clients can behave better and retry appropriately.
7) Handle quotas separately from short-term limits
A good design separates:
- short-term rate limit to prevent spikes/abuse
- monthly quota to control overall usage and billing
Example:
- 100 requests/second
- 1,000,000 requests/month
A request must pass both checks.
8) Protect against abuse patterns
Public APIs often need more than simple per-key limits:
- IP-based throttling for anonymous traffic
- CAPTCHA or proof-of-work for signups or suspicious actions
- Endpoint-specific limits for expensive operations
- Separate limits for write vs read operations
- Higher scrutiny for failed auth attempts
- Per-tenant fairness to stop one customer from starving others
9) Support tiering and overrides
Design your system so you can:
- assign different limits by plan
- grant temporary overrides
- create internal/admin bypasses
- adjust policies without redeploying code
Store policies in config/service registry rather than hardcoding them.
10) Monitor and tune
Track:
- request volumes
- 429 rates
- top consumers
- burst patterns
- false positives
- backend saturation
You’ll likely need to tune limits after observing real traffic.
11) Consider distributed consistency tradeoffs
At high scale, exact global counting can be expensive. Options:
- Strong consistency: more accurate, higher latency
- Eventual consistency: faster, small overshoot possible
- Local + global hybrid: fast local limits with periodic global reconciliation
For public APIs, many teams accept small overshoots in exchange for performance, especially for burst handling.
12) Good practical architecture
A common setup:
- Client
- CDN/WAF
- API Gateway with rate limiting
- Redis-backed distributed counters
- App services
- Billing/quota service for monthly usage
13) Example policy logic
Pseudo-flow:
if request exceeds IP abuse threshold:
reject 429
if request exceeds API key rate limit:
reject 429
if request exceeds monthly quota:
reject 429 or 402/403 depending on policy
otherwise:
allow and record usage
14) Best practices summary
- Enforce at the edge
- Use token bucket for rate limiting
- Use atomic shared counters
- Separate rate limits from quotas
- Return 429 with retry guidance
- Make policies configurable
- Monitor and adjust continuously
If you want, I can also give you:
- a Redis-based implementation pattern,
- example API gateway configs (NGINX, Kong, Envoy, AWS API Gateway), or
- a database-backed quota design for billing.