Prompt
How do I apply quotas per tenant in an API gateway?
Latest observation
To apply quotas per tenant in an API gateway, you generally need to:
- Identify the tenant on each request
- Map that tenant to a quota policy
- Enforce request limits in the gateway
- Track usage across all gateway instances
1) Identify the tenant
The gateway must know which tenant is making the request. Common ways:
- API key per tenant
- JWT claim like
tenant_id - mTLS client certificate
- Custom header such as
X-Tenant-ID(less secure unless validated)
Best practice is to derive tenant identity from a trusted auth token rather than a user-supplied header.
2) Define quotas per tenant
A quota can be:
- Requests per second/minute/hour
- Total requests per day/month
- Burst limits
- Concurrent request limits
- Bandwidth/data transfer limits
Example:
- Tenant A: 1000 req/min
- Tenant B: 100 req/min
- Tenant C: 10 req/sec burst, 5000 req/day
3) Enforce in the gateway
Most gateways support rate limiting / quota plugins or policies. The gateway checks the tenant identity, looks up the quota, and decides whether to allow or reject the request.
Typical behavior:
- If under quota → forward request
- If over quota → return
429 Too Many Requests
4) Store usage centrally
If you run multiple gateway replicas, local memory is not enough. Use a shared backend such as:
- Redis
- DynamoDB
- Cassandra
- PostgreSQL (for lower scale)
This ensures all gateway instances count against the same quota.
Implementation patterns
A. Static tenant-based policies
You preconfigure quotas by tenant ID.
Example logic:
- Read
tenant_id - Match against policy table
- Apply rate limiter key =
tenant_id
This is simple and works well when tenant count is manageable.
B. Dynamic quota lookup
The gateway queries a policy service or config store:
tenant_id -> quota planplan -> limits
This is useful when quotas change often or tenants have billing plans.
C. Hierarchical quotas
You may want both:
- Global quota for all tenants combined
- Per-tenant quota
- Per-user quota within tenant
- Per-route quota for expensive endpoints
Example:
- Tenant A: 1000 req/min total
- Tenant A
/reports: 100 req/min - Tenant A user123: 50 req/min
Example with a gateway plugin
A typical plugin flow:
- Authenticate request
- Extract
tenant_id - Build a quota key, e.g.
quota:tenant:A - Increment counter in Redis
- If counter exceeds limit, reject with
429
Pseudo-logic:
tenant = getTenantFromJWT()
limit = lookupQuota(tenant)
count = redis.increment("quota:" + tenant + ":" + currentMinute)
if count > limit:
return 429
else:
forward request
For distributed safety, use atomic Redis operations or a gateway-native rate limit service.
Important design considerations
Fairness
Decide whether quotas are:
- Fixed window: simpler, can cause bursts at window boundaries
- Sliding window: smoother, more accurate
- Token bucket/leaky bucket: best for burst handling
Multi-dimensional limits
Often you’ll want both:
- Rate limit: requests per time window
- Quota: total usage over a longer period
Fail-open vs fail-closed
If the quota backend is unavailable:
- Fail-closed: safer, but may block all traffic
- Fail-open: better availability, but risks overuse
Identity security
Never trust tenant IDs coming directly from the client unless authenticated and authorized.
Observability
Track:
- Allowed requests
- Rejected requests
- Per-tenant usage
- Limit breaches
This helps with billing and troubleshooting.
Common gateway support
Examples of how this is typically done in popular gateways:
- Kong: rate-limiting plugin, key-auth/JWT plugins
- Apigee: quota policies, spike arrest
- AWS API Gateway: usage plans + API keys, throttling
- NGINX / NGINX Plus: limit_req / limit_conn with keys
- Envoy: rate limit service integration
- Tyk: per-tenant plans and quotas
Recommended approach
For most systems:
- Authenticate with JWT or API key
- Derive
tenant_id - Use Redis-backed token bucket
- Apply both:
- per-tenant burst/rate limit
- per-tenant monthly quota
- Emit metrics per tenant
If you want, I can show:
- A generic architecture diagram
- A Kong/Apigee/AWS API Gateway example
- A Redis-based implementation pattern