Prompt
How do I set up API rate limiting by customer or app?
Latest observation
To set up API rate limiting by customer or app, you generally want to rate limit on an identifier that is stable and unique per caller, such as:
- Customer ID
- API key / client ID
- OAuth client ID
- User ID
- Account ID / organization ID
- IP address as a fallback, but not ideal for customer-level limits
Common approaches
1) Use an API key or client ID per app/customer
Best for app-level limiting.
How it works
- Each app gets its own API key or client ID.
- Your gateway or backend tracks requests per key.
- Enforce limits like:
- 100 requests/minute per app
- 1,000 requests/hour per customer
Pros
- Simple
- Easy to attribute usage
- Works well for third-party integrations
2) Use a customer/account identifier from auth context
Best for customer-level limiting.
How it works
- Authenticate the request using JWT, session, or API key.
- Extract
customer_id,tenant_id, ororg_idfrom the token or lookup. - Apply limits based on that ID.
Example
- Customer A: 500 requests/minute
- Customer B: 5,000 requests/minute
3) Combine dimensions
You can rate limit on multiple dimensions at once:
- Per customer
- Per app within customer
- Per endpoint
- Per region
- Per IP
- Global system limit
This is useful to prevent one app from exhausting the whole customer quota.
Example:
- 1,000 requests/minute per customer
- 200 requests/minute per app
- 20 requests/second per endpoint
Where to implement it
Option A: API gateway
Good if you want centralized enforcement.
Examples:
- Kong
- NGINX
- Apigee
- AWS API Gateway
- Azure API Management
- Cloudflare
- Envoy
Benefits
- Offloads work from backend
- Consistent enforcement
- Easier to manage quotas globally
Option B: Application middleware
Good if limits depend on app-specific business logic.
You can implement in:
- Express / Node.js middleware
- Django / FastAPI / Flask middleware
- Spring filters/interceptors
- Go middleware
Benefits
- Flexible
- Can use custom business rules
- Easy to tie to internal user/account models
Option C: Distributed rate limiting service
Best for high-scale systems.
Use:
- Redis
- Memcached
- DynamoDB
- PostgreSQL with careful design
- Dedicated service like Envoy ratelimit
Common pattern
- Store counters by key, like:
rate:customer:123rate:app:abc
- Use token bucket or sliding window algorithms
- Shared storage ensures limits are enforced across multiple servers
Recommended algorithm
Token bucket
Usually the best default.
How it works
- Each key has a bucket with tokens.
- Requests consume tokens.
- Tokens refill over time.
Why it’s good
- Smooth traffic bursts
- Easy to reason about
- Common in gateways and proxies
Fixed window
Simpler but less precise.
Example:
- 100 requests per minute
- Counter resets every minute
Downside
- Can allow bursts at window boundaries
Sliding window / leaky bucket
More accurate, but slightly more complex.
Practical setup steps
1) Choose the identifier
Decide what represents the caller:
- Customer
- App
- Both
Example:
tenant_idfor customer limitsclient_idfor app limits
2) Define limits
Set policies, such as:
- Free tier: 60 req/min
- Pro tier: 600 req/min
- Enterprise: 10,000 req/min
Also decide:
- Per second vs per minute vs per day
- Burst allowance
- Endpoint-specific overrides
3) Add identification to requests
Include the identifier in:
- API key metadata
- JWT claims
- Request headers from trusted auth layer
- Session lookup
Avoid trusting arbitrary client-sent headers unless they’re validated by your auth layer.
4) Store counters centrally
If you run multiple API servers, use shared state:
- Redis is common
- Use atomic increments or Lua scripts for correctness
5) Return proper HTTP responses
When limit is exceeded:
- Return
429 Too Many Requests - Include
Retry-After - Optionally include rate limit headers:
X-RateLimit-LimitX-RateLimit-RemainingX-RateLimit-Reset
6) Monitor and tune
Track:
- Per-customer usage
- Rejected requests
- Peak traffic
- Abuse patterns
- False positives
Adjust limits based on real usage.
Example policies
Per customer
- Basic: 100 req/min
- Pro: 1,000 req/min
- Enterprise: 10,000 req/min
Per app
- App A: 300 req/min
- App B: 50 req/min
Per endpoint
/login: 5 req/min per user or IP/search: 30 req/sec per user/export: 10 req/hour per account
Important security notes
- Don’t rely only on IP for customer-level limits.
- Make sure identifiers come from authenticated identity, not user-controlled input.
- Consider separate limits for:
- auth endpoints
- expensive queries
- write operations
- Add abuse protections for:
- bursts
- credential stuffing
- scraping
- noisy neighbors
Simple Redis-based example idea
Key pattern:
rl:{customer_id}:{window}
Example flow:
- On request, identify customer/app.
- Increment counter in Redis.
- If counter exceeds threshold, reject with 429.
- Set TTL so the key expires automatically.
For stronger correctness at scale, use:
- token bucket in Redis
- Lua script for atomic check/update
If you want the quickest path
If you already use an API gateway, configure rate limits there using:
- API key
- client ID
- JWT claims
- org/customer ID if supported
If not, implement middleware with Redis and token bucket logic.
If you want, I can give you:
- a Redis + Node.js example,
- a FastAPI/Django example, or
- an API gateway configuration example.