Prompt

How do I set up retries and fallbacks for an LLM API?

Artificial Intelligence / AI Platforms1 observationLast seen Oct 8, 2026

Latest observation

Oct 8, 2026OpenAI APIWeb search: off

A good retry/fallback setup for an LLM API usually has three layers:

  1. Retry transient failures
  2. Fallback to a cheaper or more reliable model/provider
  3. Degrade gracefully when all else fails

1) Retry only the failures that are worth retrying

Retry:

  • timeouts
  • rate limits (429)
  • transient server errors (500, 502, 503, 504)
  • network errors

Do not retry:

  • invalid prompts
  • auth errors
  • bad request/schema errors
  • content policy failures

Use:

  • exponential backoff
  • jitter to avoid thundering herds
  • a max retry count
  • an overall timeout budget

Example retry policy

  • Try up to 3 times
  • Backoff: 0.5s, 1s, 2s
  • Add random jitter ±20%
  • Stop if total elapsed time exceeds 10s

2) Add fallback paths

Common fallback order:

Model fallback

  • Primary: best-quality model
  • Fallback 1: cheaper/smaller model from same provider
  • Fallback 2: another provider
  • Fallback 3: rules-based or cached response

Request simplification fallback

If a request is too expensive or large:

  • shorten context
  • reduce max tokens
  • disable optional tools/function calling
  • switch to a summarization or extraction prompt

Provider fallback

If one provider is down:

  • route to another provider with a similar model
  • normalize prompts/responses across providers

3) Define your fallback decision logic

Use different policies for different failure types:

  • 429 / rate limit → retry with longer backoff, maybe switch to a lower-priority model
  • 5xx / timeout → retry same provider first, then fallback model/provider
  • context too long → automatically trim/summarize, don’t retry unchanged
  • invalid response format → one retry with a stricter prompt/schema, then fallback

4) Protect your app with guardrails

Add:

  • circuit breaker: stop calling a failing provider for a cooldown period
  • request deduplication: avoid repeated identical retries
  • idempotency keys where supported
  • observability: log attempt count, latency, error type, fallback used
  • budget limits: cap cost per request

5) A simple implementation pattern

Pseudo-logic:

def call_llm(request):
    attempts = [
        {"model": "best-model", "provider": "A"},
        {"model": "smaller-model", "provider": "A"},
        {"model": "backup-model", "provider": "B"},
    ]

    for attempt in attempts:
        for retry in range(3):
            try:
                return llm_api_call(request, **attempt)
            except RetryableError as e:
                sleep(backoff_with_jitter(retry))
                continue
            except NonRetryableError:
                break

    return fallback_response()

6) Practical recommendations

  • Keep retries short and bounded
  • Don’t retry blindly on every error
  • Fall back to a known-safe response if the LLM is unavailable
  • Prefer structured outputs and validate them before accepting them
  • Log enough data to know whether retries are helping or just increasing cost

7) If you want a production-ready setup

Typical stack:

  • retry middleware with exponential backoff
  • circuit breaker
  • multi-provider router
  • prompt/version registry
  • response validator
  • fallback cache or templated response

If you want, I can show you:

  • a Python example using tenacity
  • a Node.js example
  • a multi-provider fallback architecture
  • or a production checklist for LLM retries/failovers

Brands

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.