Prompt
api gateway rate limiting by tenant
Latest observation
Oct 6, 2026
OpenAI APIWeb search: off
To rate limit an API Gateway by tenant, the usual pattern is:
- Identify the tenant on each request
- Apply a quota/rate key per tenant
- Return 429 when a tenant exceeds its limit
- Keep limits configurable per tenant if needed
Common approaches
1) API key per tenant
Best if each tenant can have its own credential.
- Issue a unique API key to each tenant
- Configure the gateway to use the API key as the rate-limit key
- Set per-key quotas/rates
Pros
- Simple
- Native support in many gateways
- Easy reporting per tenant
Cons
- API key must be present on every request
- Not ideal if tenants share credentials or use end-user tokens
2) Tenant ID from JWT / auth token
Best when tenants authenticate with OAuth/JWT.
- Extract
tenant_idclaim from the token - Use it as the rate-limit identifier
- Enforce limits in the gateway or via a plugin/middleware
Pros
- No extra API key management
- Works well in multi-tenant SaaS
Cons
- Requires gateway support for custom claim-based rate limiting
- Needs secure token validation first
3) Header-based tenant identification
If requests carry something like X-Tenant-Id.
- Read tenant id from a trusted header
- Rate limit by that value
Important
- Only do this if the header is set by trusted clients or injected after auth
- Otherwise users can spoof tenant ids
Gateway-specific patterns
AWS API Gateway
- Use Usage Plans + API Keys for per-tenant limits
- For JWT/custom tenant claims, AWS API Gateway itself is limited; often you use:
- Lambda authorizer
- WAF rate-based rules
- Custom backend throttling
- Or move to Amazon CloudFront/WAF + backend logic
Kong
- Use
rate-limitingplugin - Key by consumer, credential, header, or custom identifier
- Tenants can map to consumers
NGINX / Ingress
- Use
limit_reqwith a key like$http_x_tenant_idor JWT claim via auth module - Usually needs custom config or Lua/OpenResty for dynamic per-tenant limits
Apigee
- Strong support for quota policies
- Can key quotas by developer app, custom attribute, or extracted tenant value
Envoy / Istio
- Use external rate limit service
- Send tenant identity as descriptor
- Rate-limit service applies per-descriptor quotas
Recommended design for SaaS multi-tenancy
Option A: Fixed tier limits
Example:
- Free: 100 req/min
- Pro: 1000 req/min
- Enterprise: custom
Flow:
- Auth succeeds
- Gateway extracts
tenant_id - Gateway calls rate-limit service with descriptor
{tenant_id, tier} - Service approves/denies
Option B: Hybrid limits
Apply:
- Per-tenant global limit
- Per-user within tenant limit
- Burst limit
- Endpoint-specific limits for expensive routes
Example:
- Tenant: 500 req/min
/search: 50 req/min/export: 5 req/min
Best practices
- Prefer tenant id from validated auth token, not user-controlled headers
- Use distributed rate limiting if gateway runs multiple instances
- Add burst + sustained limits
- Return helpful
429 Too Many RequestswithRetry-After - Log tenant id, limit, and decision for observability
- Make limits configurable without redeploying if possible
If you want to implement it quickly
Tell me which gateway you use:
- AWS API Gateway
- Kong
- NGINX
- Apigee
- Azure API Management
- Envoy/Istio
- Something else
And I can give you the exact config/example for rate limiting by tenant.