Prompt
I'm building a public API and need quotas, logs, and throttling
Latest observation
Oct 5, 2026
OpenAI APIWeb search: off
For a public API, it helps to separate three concerns:
- Quotas — how much a client is allowed to use over time
- Throttling / rate limiting — how fast they can send requests right now
- Logs / observability — how you detect abuse, debug issues, and prove usage
1) Quotas
Use quotas for longer-term limits, such as:
- Requests per day / month
- Requests per minute per API key
- Data transferred
- Number of active resources
- Concurrent jobs
- Feature-specific limits
Common patterns:
- Per API key / tenant
- Per user
- Per organization
- Per endpoint
- Per plan tier: free, pro, enterprise
Example:
- Free: 10,000 requests/month, 5 req/sec burst
- Pro: 1,000,000 requests/month, 50 req/sec burst
- Enterprise: custom
2) Throttling
Use throttling to protect your service from sudden spikes and abuse.
Typical algorithms:
- Token bucket: allows bursts, good for APIs
- Leaky bucket: smooths traffic
- Fixed window: simple, but can be unfair at boundaries
- Sliding window: more accurate, slightly more complex
A practical setup:
- Per-second/minute rate limit
- Burst allowance
- Concurrency limit for expensive operations
- Per-IP fallback limit for anonymous traffic
Example:
- 10 req/sec sustained
- Burst to 30
- 3 concurrent export jobs
3) Logs
Log at multiple levels:
Request logs
Capture:
- Timestamp
- Request ID / correlation ID
- API key or tenant ID
- Endpoint
- Method
- Status code
- Latency
- Payload size
- Rate-limit decision
- Quota remaining
Security/audit logs
Capture:
- Authentication failures
- Permission denials
- Key creation/revocation
- Quota changes
- Admin actions
- Suspicious patterns
Error logs
Capture:
- Stack traces
- Upstream failures
- Timeouts
- Retry counts
Metrics
In addition to logs, track:
- Requests per minute
- 4xx / 5xx rates
- Latency percentiles
- Throttled requests
- Quota exhaustion events
Response headers to expose limits
If this is a public API, return clear limit information:
X-RateLimit-LimitX-RateLimit-RemainingX-RateLimit-ResetRetry-Afterwhen throttled
Example 429 response:
HTTP/1.1 429 Too Many Requests
Retry-After: 60
Content-Type: application/json
{
"error": "rate_limited",
"message": "Too many requests. Try again in 60 seconds."
}
Recommended architecture
A common design:
- API gateway / edge proxy handles basic throttling
- Auth service identifies tenant/API key
- Quota service tracks monthly/daily usage
- Centralized logging/metrics for observability
- Async audit pipeline for immutable security logs
For scale, store counters in:
- Redis for real-time rate limiting
- Database / warehouse for durable quota accounting
- Log pipeline like ELK, OpenSearch, Datadog, Splunk, or Cloud Logging
Best practices
- Use multi-dimensional limits: per key, per user, per IP, per endpoint
- Differentiate 429 from 403:
- 429 = too many requests
- 403 = not allowed / plan restriction
- Make limits predictable and documented
- Support idempotency keys for retried writes
- Exempt or separately limit webhooks/internal health checks
- Protect expensive endpoints more aggressively
- Keep logs PII-minimized and retention-bounded
A simple policy example
- Anonymous: 60 req/min per IP
- Registered free user: 100 req/min, 10k/day
- Paid user: 1000 req/min, 1M/month
- Export endpoint: 1 concurrent job, 10/hour
- All clients: burst up to 2x sustained rate
If you want, I can also help you design:
- a quota schema
- a rate-limiting algorithm
- a database/Redis implementation
- or an OpenAPI/HTTP header standard for your API.