Prompt

How do I handle retries and fallback models with an LLM API?

Artificial Intelligence / AI Platforms1 observationLast seen Oct 9, 2026

Latest observation

Oct 9, 2026OpenAI APIWeb search: off

A good pattern is to treat retries and fallback models as client-side orchestration, not something you leave to chance.

1) Retry only for transient failures

Retry when the failure is likely temporary, for example:

  • HTTP 429 rate limit
  • HTTP 500, 502, 503, 504
  • network timeouts / connection resets

Don’t blindly retry for:

  • invalid API key
  • bad request / malformed prompt
  • context too long
  • content policy rejections

2) Use exponential backoff with jitter

Avoid retrying immediately in a tight loop. Instead:

  • wait a short time
  • then increase delay on each retry
  • add jitter so many clients don’t retry at once

Typical pattern:

  • attempt 1: no wait
  • attempt 2: 200–500 ms
  • attempt 3: 1–2 s
  • attempt 4: 4–8 s

Cap the maximum delay and the total retry time.

3) Keep retries bounded

Set clear limits:

  • max attempts: e.g. 3–5
  • max total elapsed time
  • idempotency or deduplication if your workflow can cause duplicates

For generation requests, retries are usually safe, but if the call triggers side effects downstream, ensure the full workflow is idempotent.

4) Add fallback models by priority

If your primary model fails after retries, fall back to a cheaper or more available model:

Example order:

  1. gpt-4.1
  2. gpt-4.1-mini
  3. gpt-4o-mini

Use fallback for:

  • repeated transient failures
  • overload
  • cost control
  • latency constraints

Be careful: fallback models may have different quality, formatting, or tool-use behavior.

5) Normalize prompts across models

To make fallback safe:

  • keep system instructions consistent
  • avoid model-specific prompt tricks
  • request structured output when possible
  • validate output before using it

If the fallback model returns malformed output, either:

  • repair/parse it
  • or escalate to a human / alternate workflow

6) Separate failure types

A practical policy:

  • Retry same model for transient errors
  • Fallback to another model after retry budget is exhausted
  • Fail fast on permanent errors

7) Log enough to debug

Record:

  • model used
  • attempt count
  • error type/status code
  • latency
  • whether fallback was used
  • final response quality metrics if you have them

This helps you tune retry thresholds and fallback ordering.

8) Example strategy

A simple policy:

  • Try primary model up to 3 times
  • Retry on 429, 5xx, timeouts
  • Backoff: 0.5s, 1s, 2s with jitter
  • If still failing, switch to fallback model and try 2 times
  • If all fail, return a controlled error

9) Python-style pseudocode

import random
import time

RETRYABLE_STATUSES = {429, 500, 502, 503, 504}

def call_model(model, prompt):
    # replace with actual API call
    pass

def should_retry(error):
    status = getattr(error, "status_code", None)
    return status in RETRYABLE_STATUSES or isinstance(error, TimeoutError)

def run_with_fallback(prompt):
    models = [
        ("primary-model", 3),
        ("fallback-model", 2),
    ]

    for model, max_attempts in models:
        for attempt in range(1, max_attempts + 1):
            try:
                return call_model(model, prompt)
            except Exception as e:
                if not should_retry(e):
                    raise

                if attempt == max_attempts:
                    break

                base_delay = 0.5 * (2 ** (attempt - 1))
                jitter = random.uniform(0, base_delay * 0.2)
                time.sleep(base_delay + jitter)

    raise RuntimeError("All models failed")

10) Best practice summary

  • Retry transient failures only
  • Use exponential backoff + jitter
  • Cap retries
  • Fall back to another model after retry exhaustion
  • Validate outputs from fallback models
  • Log and monitor everything

If you want, I can also give you:

  • a production-ready Python implementation
  • a Node.js example
  • or a retry/fallback policy template for your specific API stack

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.