Prompt
api gateway rate limit by tenant
Latest observation
To rate limit an API Gateway by tenant, you usually need a way to identify the tenant on each request, then apply limits at that identity level.
Common approaches
1) Use a tenant ID in a header or claim
Typical options:
X-Tenant-Idheader- JWT/OAuth claim like
tenant_id - API key mapped to a tenant
Then enforce:
- requests per second/minute
- burst limits
- daily quotas
- concurrency limits
Implementation patterns
A. Gateway-native rate limiting
If your API gateway supports custom rate limits, configure it to key off tenant identity.
Examples:
- Kong: rate-limiting plugin with custom consumer/credential mapping
- Apigee: quota/rate limit policies using variables like app, developer, or custom attributes
- AWS API Gateway: usage plans + API keys, or Lambda authorizer/custom logic
- NGINX / NGINX Plus:
limit_reqwith variables - Envoy: local/global rate limit service with descriptors
- Azure API Management:
rate-limit-by-key
This is often the best approach if your gateway supports it.
B. External rate limit service
If tenant-based limits are dynamic or complex, use a centralized store/service:
- Redis
- Memcached
- DynamoDB
- custom rate limit service
Flow:
- Gateway extracts tenant ID
- Gateway calls rate limit service
- Service checks quota for that tenant
- Allow or deny
This works well for:
- per-tenant plans
- per-tenant burst + sustained limits
- overrides and temporary promotions
- distributed gateways
C. API key per tenant
Issue each tenant a unique API key or client credential. Then rate limit by key.
Pros:
- simple
- clean tenant mapping
- easy quota management
Cons:
- tenant may need multiple keys for different apps
- less flexible if one tenant has many users or services
Key design considerations
Identity source
Prefer a trusted identity source:
- JWT claim from authenticated token
- mTLS client cert mapped to tenant
- API key tied to tenant
Avoid trusting a plain header like X-Tenant-Id unless the gateway or auth layer validates it.
Shared tenant limits
Decide whether limits apply to:
- entire tenant across all users/apps
- per end user within tenant
- per tenant per API
- per tenant per region/environment
Example:
- tenant A: 1000 req/min total
- tenant A on
/reports: 100 req/min - tenant A burst: 50 req/sec
Fairness and burst handling
Use:
- token bucket for burstable traffic
- leaky bucket for steady smoothing
- quotas for daily/monthly caps
Multi-gateway deployment
If you run multiple gateway instances, local in-memory counters won’t be enough. Use:
- centralized Redis
- distributed rate limit service
- gateway-native global rate limiting
Example logic
Pseudo-flow:
Authenticate request
Extract tenant_id
Look up tenant plan
Apply rate limit:
if requests_this_window >= tenant_limit:
return 429 Too Many Requests
else:
forward request
Example policies
Per-tenant quota
- Bronze: 100 req/min
- Silver: 1000 req/min
- Gold: 10000 req/min
Per-tenant + per-endpoint
/search: 50 req/sec/export: 5 req/min/admin: 10 req/min
Recommended practice
Best practice is usually:
- Authenticate the caller
- Derive tenant ID from trusted auth context
- Apply rate limit using tenant ID
- Use Redis or gateway-native global limit if running multiple instances
- Return 429 with retry-after when exceeded
If you want, I can show you how to do this in a specific gateway like Kong, Apigee, AWS API Gateway, NGINX, or Azure APIM.