Prompt

How do I set up retries and timeouts in an API gateway?

Technology / API Platforms1 observationLast seen Oct 6, 2026

Latest observation

Oct 6, 2026OpenAI APIWeb search: off

To set up retries and timeouts in an API gateway, you usually configure them at the route/service/upstream level rather than in each client. The exact syntax depends on the gateway (Kong, Apigee, AWS API Gateway, NGINX, Envoy, Azure API Management, etc.), but the concepts are the same.

1) Decide what should time out

There are usually a few different timeout layers:

  • Client timeout: how long the caller waits
  • Gateway timeout: how long the gateway waits for the upstream
  • Upstream/service timeout: how long the backend is allowed to take

A common pattern is:

  • Client timeout slightly higher than gateway timeout
  • Gateway timeout slightly higher than backend’s normal response time
  • Upstream services also enforce their own internal timeouts

2) Configure request timeout

Set a maximum duration for the gateway to wait for the backend before failing the request.

Typical settings:

  • Connect timeout: time to establish TCP connection
  • Read/response timeout: time waiting for response data
  • Write timeout: time allowed to send the request body
  • Overall timeout: total request duration

Example ideas:

  • Connect timeout: 1s
  • Read timeout: 5s
  • Total timeout: 6s

3) Configure retries

Retries help with transient failures like:

  • network timeouts
  • connection resets
  • upstream 502/503/504 errors
  • intermittent DNS issues

Good retry controls include:

  • Retry count: e.g. 2 or 3 retries
  • Retry conditions:
    • connection failure
    • timeout
    • specific 5xx status codes
  • Backoff:
    • fixed delay or exponential backoff
  • Jitter:
    • randomization to avoid retry storms

4) Be careful with non-idempotent requests

Only retry safely for requests that won’t cause duplicate side effects, such as:

  • GET
  • HEAD
  • PUT/DELETE if your API is designed to be idempotent

Be very cautious with:

  • POST
  • payment, order creation, or any write operation

If you must retry writes, use:

  • idempotency keys
  • deduplication on the backend

5) Prefer retrying only on transient failures

Do not retry on:

  • validation errors (400)
  • auth failures (401, 403)
  • not found (404)
  • permanent backend errors

Retry on:

  • 502 Bad Gateway
  • 503 Service Unavailable
  • 504 Gateway Timeout
  • connection failures
  • read timeouts

6) Add circuit breaking and rate limiting

Retries can amplify outages. To avoid that:

  • use a circuit breaker to stop sending traffic to a failing service
  • set max retries low
  • use rate limiting
  • monitor error rates and latency

7) Example configurations

NGINX-style example

location /api/ {
    proxy_connect_timeout 1s;
    proxy_send_timeout 5s;
    proxy_read_timeout 5s;
    proxy_next_upstream error timeout http_502 http_503 http_504;
    proxy_next_upstream_tries 3;
}

Envoy-style example

retry_policy:
  retry_on: "5xx,connect-failure,refused-stream,reset,unavailable,cancelled"
  num_retries: 2
  per_try_timeout: 2s

timeout: 6s

Kong-style example

Kong often uses plugins and upstream/service settings depending on version and deployment, with:

  • service connect_timeout
  • service read_timeout
  • service write_timeout
  • retries on upstream connection issues via proxy behavior or plugins

8) Test and tune

After configuring:

  • simulate slow backend responses
  • simulate 502/503/504 responses
  • verify retry count and latency
  • ensure you are not causing duplicate writes
  • inspect logs and metrics

Practical recommendation

A safe starting point for many APIs:

  • 1–2 retries max
  • exponential backoff with jitter
  • retry only on transient 5xx/timeouts
  • timeout slightly above normal p95/p99 latency
  • no retries for non-idempotent operations unless using idempotency keys

If you tell me which API gateway you’re using, I can give you the exact configuration example.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.