Prompt
how to set api quotas per customer
Latest observation
To set API quotas per customer, you typically need identity, metering, enforcement, and reporting.
1) Identify each customer
Give every customer a unique identifier, such as:
customer_id- API key tied to a customer
- OAuth client ID / tenant ID
- account/org ID
This ID is what you’ll use to look up their quota.
2) Define quota dimensions
Decide what you want to limit. Common quota types:
- Requests per minute/hour/day/month
- Concurrent requests
- Bandwidth / payload size
- Compute units / tokens / credits
- Per-endpoint limits (e.g.,
/searchgets 1000/day)
Example:
- Free: 1000 requests/day
- Pro: 100,000 requests/day
- Enterprise: custom
3) Store plan and limits
Keep customer quota settings in a database or config store:
Example fields:
customer_idplanlimit_requests_per_minlimit_requests_per_daylimit_burstreset_periodoverage_policy
4) Enforce quotas at the API gateway or middleware
Check quota on every request before processing.
Typical flow:
- Authenticate request
- Resolve
customer_id - Read current usage
- Compare against limit
- Allow or reject request
If over quota, return:
- HTTP
429 Too Many Requests - Include
Retry-Afterheader if relevant
5) Track usage efficiently
Use a fast store like:
- Redis
- Memcached
- Distributed counters
- API gateway built-in rate limiting
For example, Redis can maintain counters per customer:
usage:{customer_id}:dayusage:{customer_id}:minute
Use atomic increments and expirations so counters reset automatically.
6) Support burst and smoothing
Often you want both:
- Hard quota: max per day/month
- Rate limit: max per second/minute
This prevents customers from exhausting all quota instantly.
Token bucket or leaky bucket algorithms are commonly used.
7) Add observability and alerts
Track:
- Current usage
- Remaining quota
- Quota violations
- Top customers by usage
Alert customers when they reach:
- 80%
- 90%
- 100%
8) Offer customer-visible usage APIs
Provide endpoints like:
GET /usageGET /quota
Return:
- limit
- consumed
- remaining
- reset time
9) Handle multi-tenant and enterprise cases
For enterprise customers, quotas may be:
- shared across teams
- per workspace
- per product
- custom negotiated
Model it with hierarchy:
- org quota
- project quota
- user quota
10) Example implementation idea
Pseudo-logic:
def handle_request(request):
customer_id = authenticate(request)
quota = get_quota(customer_id)
usage = increment_usage_atomic(customer_id)
if usage > quota.daily_limit:
return 429, {"error": "quota exceeded"}
return process_request(request)
Common best practices
- Use server-side enforcement; don’t trust clients
- Make counters atomic
- Support rolling windows if exact calendar resets aren’t needed
- Separate rate limiting from billing quota
- Log all quota decisions for auditing
If you want, I can show:
- a Redis-based quota design,
- an API gateway approach (Kong / NGINX / Apigee), or
- a database schema for per-customer quotas.