Prompt

How do I apply quotas per tenant in an API gateway?

Technology / API Platforms1 observationLast seen Oct 6, 2026

Latest observation

Oct 6, 2026OpenAI APIWeb search: off

To apply quotas per tenant in an API gateway, you generally need to:

  1. Identify the tenant on each request
  2. Map that tenant to a quota policy
  3. Enforce request limits in the gateway
  4. Track usage across all gateway instances

1) Identify the tenant

The gateway must know which tenant is making the request. Common ways:

  • API key per tenant
  • JWT claim like tenant_id
  • mTLS client certificate
  • Custom header such as X-Tenant-ID (less secure unless validated)

Best practice is to derive tenant identity from a trusted auth token rather than a user-supplied header.

2) Define quotas per tenant

A quota can be:

  • Requests per second/minute/hour
  • Total requests per day/month
  • Burst limits
  • Concurrent request limits
  • Bandwidth/data transfer limits

Example:

  • Tenant A: 1000 req/min
  • Tenant B: 100 req/min
  • Tenant C: 10 req/sec burst, 5000 req/day

3) Enforce in the gateway

Most gateways support rate limiting / quota plugins or policies. The gateway checks the tenant identity, looks up the quota, and decides whether to allow or reject the request.

Typical behavior:

  • If under quota → forward request
  • If over quota → return 429 Too Many Requests

4) Store usage centrally

If you run multiple gateway replicas, local memory is not enough. Use a shared backend such as:

  • Redis
  • DynamoDB
  • Cassandra
  • PostgreSQL (for lower scale)

This ensures all gateway instances count against the same quota.


Implementation patterns

A. Static tenant-based policies

You preconfigure quotas by tenant ID.

Example logic:

  • Read tenant_id
  • Match against policy table
  • Apply rate limiter key = tenant_id

This is simple and works well when tenant count is manageable.

B. Dynamic quota lookup

The gateway queries a policy service or config store:

  • tenant_id -> quota plan
  • plan -> limits

This is useful when quotas change often or tenants have billing plans.

C. Hierarchical quotas

You may want both:

  • Global quota for all tenants combined
  • Per-tenant quota
  • Per-user quota within tenant
  • Per-route quota for expensive endpoints

Example:

  • Tenant A: 1000 req/min total
  • Tenant A /reports: 100 req/min
  • Tenant A user123: 50 req/min

Example with a gateway plugin

A typical plugin flow:

  1. Authenticate request
  2. Extract tenant_id
  3. Build a quota key, e.g. quota:tenant:A
  4. Increment counter in Redis
  5. If counter exceeds limit, reject with 429

Pseudo-logic:

tenant = getTenantFromJWT()
limit = lookupQuota(tenant)
count = redis.increment("quota:" + tenant + ":" + currentMinute)

if count > limit:
    return 429
else:
    forward request

For distributed safety, use atomic Redis operations or a gateway-native rate limit service.


Important design considerations

Fairness

Decide whether quotas are:

  • Fixed window: simpler, can cause bursts at window boundaries
  • Sliding window: smoother, more accurate
  • Token bucket/leaky bucket: best for burst handling

Multi-dimensional limits

Often you’ll want both:

  • Rate limit: requests per time window
  • Quota: total usage over a longer period

Fail-open vs fail-closed

If the quota backend is unavailable:

  • Fail-closed: safer, but may block all traffic
  • Fail-open: better availability, but risks overuse

Identity security

Never trust tenant IDs coming directly from the client unless authenticated and authorized.

Observability

Track:

  • Allowed requests
  • Rejected requests
  • Per-tenant usage
  • Limit breaches

This helps with billing and troubleshooting.


Common gateway support

Examples of how this is typically done in popular gateways:

  • Kong: rate-limiting plugin, key-auth/JWT plugins
  • Apigee: quota policies, spike arrest
  • AWS API Gateway: usage plans + API keys, throttling
  • NGINX / NGINX Plus: limit_req / limit_conn with keys
  • Envoy: rate limit service integration
  • Tyk: per-tenant plans and quotas

Recommended approach

For most systems:

  • Authenticate with JWT or API key
  • Derive tenant_id
  • Use Redis-backed token bucket
  • Apply both:
    • per-tenant burst/rate limit
    • per-tenant monthly quota
  • Emit metrics per tenant

If you want, I can show:

  1. A generic architecture diagram
  2. A Kong/Apigee/AWS API Gateway example
  3. A Redis-based implementation pattern

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.