Prompt

How do I set up fallback between multiple LLM providers if one is slow or down?

Artificial Intelligence / AI Infrastructure1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

A good fallback setup for multiple LLM providers usually has 4 parts:

  1. A common interface
  2. Health/latency tracking
  3. A routing policy
  4. Retries + timeouts + circuit breakers

1) Normalize all providers behind one interface

Wrap each provider so your app calls them the same way:

class LLMProvider:
    def generate(self, prompt, **kwargs):
        raise NotImplementedError

Then implement one adapter per provider:

class OpenAIProvider(LLMProvider):
    def generate(self, prompt, **kwargs):
        ...

class AnthropicProvider(LLMProvider):
    def generate(self, prompt, **kwargs):
        ...

class GeminiProvider(LLMProvider):
    def generate(self, prompt, **kwargs):
        ...

This makes failover logic provider-agnostic.


2) Set strict timeouts and retries

Don’t wait forever on a slow provider.

  • Connect timeout: e.g. 2–5s
  • Request timeout: e.g. 15–30s depending on use case
  • Retries: small number, only for transient errors
  • Backoff: exponential with jitter

Important:

  • Retry timeouts, 429s, 5xx
  • Don’t retry most 4xx errors like bad auth or invalid input

Example policy:

  • Try provider A
  • If A times out or fails transiently, try B
  • If B fails, try C
  • Stop after N providers or return an error

3) Track provider health

Maintain live stats per provider:

  • success rate
  • p50 / p95 latency
  • timeout rate
  • error rate
  • last failure time
  • circuit breaker state

Use this to decide which provider to try first.

A simple rule:

  • Prefer the fastest healthy provider
  • Demote providers with high error/timeout rates
  • Temporarily disable a provider after repeated failures

4) Use a circuit breaker

If a provider is repeatedly failing, stop sending traffic to it for a short cooldown.

Typical states:

  • Closed: normal
  • Open: do not send requests
  • Half-open: test with a small number of requests

This prevents wasting time on a broken provider.


Recommended request flow

  1. Build a ranked provider list based on:
    • health
    • latency
    • cost
    • model capability
  2. Send request to the top provider with a timeout
  3. If it fails or is too slow:
    • log the failure
    • mark provider health
    • move to next provider
  4. Return first successful response

Example fallback pseudocode

def generate_with_fallback(prompt, providers):
    last_error = None

    for provider in providers:
        try:
            return provider.generate(prompt, timeout=20)
        except TimeoutError as e:
            last_error = e
            continue
        except TransientProviderError as e:
            last_error = e
            continue
        except PermanentProviderError:
            # Usually don't retry on permanent config/auth errors
            raise

    raise RuntimeError(f"All providers failed: {last_error}")

Better: rank by health before trying

def rank_providers(providers, health):
    return sorted(
        providers,
        key=lambda p: (
            not health[p.name].is_healthy(),
            health[p.name].p95_latency,
            health[p.name].error_rate
        )
    )

Then try them in ranked order.


Best practices

Fail open vs fail closed

Decide what to do if all providers fail:

  • Fail open: return partial response, cached answer, or simplified mode
  • Fail closed: return an error immediately

Separate by task

Use different fallback policies for:

  • chat
  • embeddings
  • tool/function calling
  • long-context summarization

A provider that is good for chat may not be good for tool calling.

Keep prompts compatible

Different providers have different formatting requirements:

  • system messages
  • tool schemas
  • token limits
  • sampling params

Normalize prompt construction in one place.

Watch out for response differences

Fallback may change:

  • style
  • safety behavior
  • function-calling format
  • JSON strictness

If you need structured output, enforce validation and re-ask if needed.


If you want production-grade reliability

Use:

  • timeouts
  • retries
  • circuit breaker
  • health scoring
  • request hedging for very latency-sensitive requests

Hedging

Instead of waiting for a slow provider to fail, you can:

  • send to provider A
  • after a short delay, send to B as backup
  • take the first valid response This reduces tail latency but costs more.

Simple architecture

Client → Router → Provider Adapters → Monitoring

  • Router picks provider
  • Adapters handle API differences
  • Monitoring feeds latency/error metrics back into routing

Minimal practical setup

If you want the simplest robust version:

  • Use 2–3 providers
  • Add a 10–20s request timeout
  • Retry once on transient errors
  • If the request fails, try the next provider
  • Track success/latency in Redis or a DB
  • Add a circuit breaker after 3–5 consecutive failures

If you want, I can give you:

  1. a Python implementation of an LLM fallback router, or
  2. a LangChain / LiteLLM / OpenRouter-based setup.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.