Prompt
Should I put rate limits at the gateway or in the app?
Latest observation
Oct 5, 2026
OpenAI APIWeb search: off
Usually: do both, but for different reasons.
Put rate limiting at the gateway when you want:
- Protection from abuse/DDoS-ish traffic
- A single enforcement point for all services
- Early rejection before requests consume app resources
- Consistent org-wide quotas across APIs
This is the right place for things like:
- per-IP limits
- per-client/API key limits
- global burst control
- shielding downstream services
Put rate limiting in the app when you need:
- Business-aware limits that depend on user/account state
- Per-endpoint or per-action rules that vary by logic
- Fine-grained control tied to permissions, plan tier, or resource cost
- Correctness for sensitive operations, even if traffic bypasses the gateway
Examples:
- 5 password reset attempts per hour per user
- different limits for free vs paid plans
- stricter limits on expensive report generation than on reads
Best practice
- Gateway: coarse, outer layer protection
- App: precise, business-level enforcement
If you only choose one:
- Choose gateway for infrastructure protection and simplicity
- Choose app for correctness of business rules
One important caveat
If you run multiple app instances, app-level rate limiting must use shared state (Redis, database, etc.) or it will be inconsistent.
Practical recommendation
Start with:
- Gateway rate limits for basic protection
- App-side limits only for rules that affect product behavior or abuse-sensitive flows
If you want, I can suggest a concrete setup for your stack (e.g. NGINX, API Gateway, Kubernetes ingress, Express, Spring, etc.).