Prompt

How do I set up API rate limiting by customer or app?

Technology / API Platforms1 observationLast seen Oct 5, 2026

Latest observation

Oct 5, 2026OpenAI APIWeb search: off

To set up API rate limiting by customer or app, you generally want to rate limit on an identifier that is stable and unique per caller, such as:

  • Customer ID
  • API key / client ID
  • OAuth client ID
  • User ID
  • Account ID / organization ID
  • IP address as a fallback, but not ideal for customer-level limits

Common approaches

1) Use an API key or client ID per app/customer

Best for app-level limiting.

How it works

  • Each app gets its own API key or client ID.
  • Your gateway or backend tracks requests per key.
  • Enforce limits like:
    • 100 requests/minute per app
    • 1,000 requests/hour per customer

Pros

  • Simple
  • Easy to attribute usage
  • Works well for third-party integrations

2) Use a customer/account identifier from auth context

Best for customer-level limiting.

How it works

  • Authenticate the request using JWT, session, or API key.
  • Extract customer_id, tenant_id, or org_id from the token or lookup.
  • Apply limits based on that ID.

Example

  • Customer A: 500 requests/minute
  • Customer B: 5,000 requests/minute

3) Combine dimensions

You can rate limit on multiple dimensions at once:

  • Per customer
  • Per app within customer
  • Per endpoint
  • Per region
  • Per IP
  • Global system limit

This is useful to prevent one app from exhausting the whole customer quota.

Example:

  • 1,000 requests/minute per customer
  • 200 requests/minute per app
  • 20 requests/second per endpoint

Where to implement it

Option A: API gateway

Good if you want centralized enforcement.

Examples:

  • Kong
  • NGINX
  • Apigee
  • AWS API Gateway
  • Azure API Management
  • Cloudflare
  • Envoy

Benefits

  • Offloads work from backend
  • Consistent enforcement
  • Easier to manage quotas globally

Option B: Application middleware

Good if limits depend on app-specific business logic.

You can implement in:

  • Express / Node.js middleware
  • Django / FastAPI / Flask middleware
  • Spring filters/interceptors
  • Go middleware

Benefits

  • Flexible
  • Can use custom business rules
  • Easy to tie to internal user/account models

Option C: Distributed rate limiting service

Best for high-scale systems.

Use:

  • Redis
  • Memcached
  • DynamoDB
  • PostgreSQL with careful design
  • Dedicated service like Envoy ratelimit

Common pattern

  • Store counters by key, like:
    • rate:customer:123
    • rate:app:abc
  • Use token bucket or sliding window algorithms
  • Shared storage ensures limits are enforced across multiple servers

Recommended algorithm

Token bucket

Usually the best default.

How it works

  • Each key has a bucket with tokens.
  • Requests consume tokens.
  • Tokens refill over time.

Why it’s good

  • Smooth traffic bursts
  • Easy to reason about
  • Common in gateways and proxies

Fixed window

Simpler but less precise.

Example:

  • 100 requests per minute
  • Counter resets every minute

Downside

  • Can allow bursts at window boundaries

Sliding window / leaky bucket

More accurate, but slightly more complex.


Practical setup steps

1) Choose the identifier

Decide what represents the caller:

  • Customer
  • App
  • Both

Example:

  • tenant_id for customer limits
  • client_id for app limits

2) Define limits

Set policies, such as:

  • Free tier: 60 req/min
  • Pro tier: 600 req/min
  • Enterprise: 10,000 req/min

Also decide:

  • Per second vs per minute vs per day
  • Burst allowance
  • Endpoint-specific overrides

3) Add identification to requests

Include the identifier in:

  • API key metadata
  • JWT claims
  • Request headers from trusted auth layer
  • Session lookup

Avoid trusting arbitrary client-sent headers unless they’re validated by your auth layer.


4) Store counters centrally

If you run multiple API servers, use shared state:

  • Redis is common
  • Use atomic increments or Lua scripts for correctness

5) Return proper HTTP responses

When limit is exceeded:

  • Return 429 Too Many Requests
  • Include Retry-After
  • Optionally include rate limit headers:
    • X-RateLimit-Limit
    • X-RateLimit-Remaining
    • X-RateLimit-Reset

6) Monitor and tune

Track:

  • Per-customer usage
  • Rejected requests
  • Peak traffic
  • Abuse patterns
  • False positives

Adjust limits based on real usage.


Example policies

Per customer

  • Basic: 100 req/min
  • Pro: 1,000 req/min
  • Enterprise: 10,000 req/min

Per app

  • App A: 300 req/min
  • App B: 50 req/min

Per endpoint

  • /login: 5 req/min per user or IP
  • /search: 30 req/sec per user
  • /export: 10 req/hour per account

Important security notes

  • Don’t rely only on IP for customer-level limits.
  • Make sure identifiers come from authenticated identity, not user-controlled input.
  • Consider separate limits for:
    • auth endpoints
    • expensive queries
    • write operations
  • Add abuse protections for:
    • bursts
    • credential stuffing
    • scraping
    • noisy neighbors

Simple Redis-based example idea

Key pattern:

  • rl:{customer_id}:{window}

Example flow:

  1. On request, identify customer/app.
  2. Increment counter in Redis.
  3. If counter exceeds threshold, reject with 429.
  4. Set TTL so the key expires automatically.

For stronger correctness at scale, use:

  • token bucket in Redis
  • Lua script for atomic check/update

If you want the quickest path

If you already use an API gateway, configure rate limits there using:

  • API key
  • client ID
  • JWT claims
  • org/customer ID if supported

If not, implement middleware with Redis and token bucket logic.

If you want, I can give you:

  1. a Redis + Node.js example,
  2. a FastAPI/Django example, or
  3. an API gateway configuration example.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.