Prompt
How do I enforce API rate limits and quotas for different customers?
Latest observation
To enforce API rate limits and quotas for different customers, use a combination of identity, policy, metering, and enforcement.
1) Identify each customer
You need a reliable way to know who is calling your API:
- API keys for machine-to-machine access
- OAuth client IDs / access tokens
- JWT claims like
sub,tenant_id, orplan - mTLS certificates for high-security environments
Every request should map to a customer ID or tenant ID.
2) Define limits and quotas per customer
Separate the concepts:
- Rate limit: how many requests are allowed in a short period
Example:100 requests/minute - Quota: how much usage is allowed over a longer period
Example:1,000,000 requests/month
You can assign different tiers:
- Free: 10 req/sec, 10k/month
- Pro: 100 req/sec, 1M/month
- Enterprise: custom limits
3) Enforce at the edge or gateway
Best place is usually an API gateway or reverse proxy:
- Kong
- Apigee
- AWS API Gateway
- Azure API Management
- NGINX / Envoy / HAProxy
- Cloudflare API Gateway
These tools often support:
- Per-key or per-consumer limits
- Burst control
- Sliding/fixed window policies
- Global or distributed counters
4) Use a counter store for distributed enforcement
If you run multiple API servers, don’t keep counters only in memory. Use shared storage:
- Redis is the most common choice
- DynamoDB, Cassandra, or a managed rate-limit service also work
Typical approach:
- On each request, increment the customer’s counter
- Check against the configured threshold
- Reject with
429 Too Many Requestsif exceeded
5) Choose a rate-limiting algorithm
Common algorithms:
- Token bucket: best for allowing bursts while controlling average rate
- Leaky bucket: smooths traffic at a constant rate
- Fixed window: simplest, but can allow burstiness at window boundaries
- Sliding window: more accurate, slightly more complex
Token bucket is often the best default.
6) Return proper headers and errors
When rejecting or nearing limits:
- Return HTTP 429
- Include headers like:
Retry-AfterX-RateLimit-LimitX-RateLimit-RemainingX-RateLimit-Reset
This helps customers back off gracefully.
7) Add per-endpoint and per-resource limits
Different endpoints often need different policies:
GET /search: strict rate limitPOST /orders: tighter controlsGET /status: more lenient- expensive endpoints: separate quota bucket
You can define limits by:
- customer
- plan
- endpoint
- method
- region
- IP address, if needed
8) Consider fairness and abuse controls
To prevent one customer from monopolizing capacity:
- Limit per customer
- Limit per IP and per API key
- Separate read vs write budgets
- Use concurrency limits for expensive operations
- Add burst limits in addition to sustained limits
9) Make quota resets explicit
For monthly or daily quotas:
- Reset on a schedule
- Or use rolling windows
- Communicate reset times clearly in docs and headers
10) Monitor and audit
Track:
- Requests per customer
- Rejections
- Spikes and anomalies
- Top endpoints by usage
This helps with:
- billing
- abuse detection
- plan tuning
- support/debugging
Practical implementation pattern
A common architecture looks like this:
- Request arrives at API gateway
- Gateway authenticates request
- Gateway extracts
customer_id - Gateway checks rate-limit/quota policy in Redis or local distributed limiter
- If allowed, forward to API
- If not allowed, return
429
Example policy model
You might store policies like:
{
"customer_id": "cust_123",
"plan": "pro",
"limits": {
"per_second": 50,
"per_minute": 1000,
"per_month": 500000
}
}
Recommendation
If you’re building this from scratch:
- Use API keys/JWTs for identity
- Put enforcement in an API gateway
- Use Redis + token bucket for distributed rate limiting
- Track monthly quotas separately from short-term rate limits
- Return 429 with standard headers
If you want, I can also give you:
- a Redis-based implementation
- an example for Node.js, Python, Go, or Java
- a gateway configuration for Kong/NGINX/AWS API Gateway