Prompt
How do I set up retries and timeouts in an API gateway?
Latest observation
To set up retries and timeouts in an API gateway, you usually configure them at the route/service/upstream level rather than in each client. The exact syntax depends on the gateway (Kong, Apigee, AWS API Gateway, NGINX, Envoy, Azure API Management, etc.), but the concepts are the same.
1) Decide what should time out
There are usually a few different timeout layers:
- Client timeout: how long the caller waits
- Gateway timeout: how long the gateway waits for the upstream
- Upstream/service timeout: how long the backend is allowed to take
A common pattern is:
- Client timeout slightly higher than gateway timeout
- Gateway timeout slightly higher than backend’s normal response time
- Upstream services also enforce their own internal timeouts
2) Configure request timeout
Set a maximum duration for the gateway to wait for the backend before failing the request.
Typical settings:
- Connect timeout: time to establish TCP connection
- Read/response timeout: time waiting for response data
- Write timeout: time allowed to send the request body
- Overall timeout: total request duration
Example ideas:
- Connect timeout:
1s - Read timeout:
5s - Total timeout:
6s
3) Configure retries
Retries help with transient failures like:
- network timeouts
- connection resets
- upstream 502/503/504 errors
- intermittent DNS issues
Good retry controls include:
- Retry count: e.g. 2 or 3 retries
- Retry conditions:
- connection failure
- timeout
- specific 5xx status codes
- Backoff:
- fixed delay or exponential backoff
- Jitter:
- randomization to avoid retry storms
4) Be careful with non-idempotent requests
Only retry safely for requests that won’t cause duplicate side effects, such as:
GETHEADPUT/DELETEif your API is designed to be idempotent
Be very cautious with:
POST- payment, order creation, or any write operation
If you must retry writes, use:
- idempotency keys
- deduplication on the backend
5) Prefer retrying only on transient failures
Do not retry on:
- validation errors (
400) - auth failures (
401,403) - not found (
404) - permanent backend errors
Retry on:
502 Bad Gateway503 Service Unavailable504 Gateway Timeout- connection failures
- read timeouts
6) Add circuit breaking and rate limiting
Retries can amplify outages. To avoid that:
- use a circuit breaker to stop sending traffic to a failing service
- set max retries low
- use rate limiting
- monitor error rates and latency
7) Example configurations
NGINX-style example
location /api/ {
proxy_connect_timeout 1s;
proxy_send_timeout 5s;
proxy_read_timeout 5s;
proxy_next_upstream error timeout http_502 http_503 http_504;
proxy_next_upstream_tries 3;
}
Envoy-style example
retry_policy:
retry_on: "5xx,connect-failure,refused-stream,reset,unavailable,cancelled"
num_retries: 2
per_try_timeout: 2s
timeout: 6s
Kong-style example
Kong often uses plugins and upstream/service settings depending on version and deployment, with:
- service
connect_timeout - service
read_timeout - service
write_timeout - retries on upstream connection issues via proxy behavior or plugins
8) Test and tune
After configuring:
- simulate slow backend responses
- simulate 502/503/504 responses
- verify retry count and latency
- ensure you are not causing duplicate writes
- inspect logs and metrics
Practical recommendation
A safe starting point for many APIs:
- 1–2 retries max
- exponential backoff with jitter
- retry only on transient 5xx/timeouts
- timeout slightly above normal p95/p99 latency
- no retries for non-idempotent operations unless using idempotency keys
If you tell me which API gateway you’re using, I can give you the exact configuration example.