Prompt

How do I set fallback rules for LLM requests when a provider fails?

Artificial Intelligence / AI Infrastructure1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

A good way to handle LLM provider failures is to define a fallback policy with clear priorities, retries, and routing rules. Here’s a practical pattern you can use.

1) Decide what counts as a failure

Common failure types:

  • Timeout
  • Rate limit / 429
  • 5xx provider errors
  • Network errors
  • Malformed response / schema violation
  • Safety refusal (depending on your use case)

Not every failure should trigger the same fallback. For example:

  • Retry the same provider for transient errors
  • Switch providers for persistent or capacity-related errors
  • Return a cached or lower-cost model for non-critical requests

2) Define a fallback chain

A simple chain might look like:

  1. Try primary provider/model
  2. If timeout/5xx/429, retry once with backoff
  3. If still failing, try secondary provider/model
  4. If still failing, try a cheaper or smaller model
  5. If everything fails, return a graceful degradation message or queue for later

Example policy:

primary: openai/gpt-4.1
fallback_1: anthropic/claude-3.5-sonnet
fallback_2: openai/gpt-4o-mini
final: cached answer / apology / manual review

3) Use error-specific rules

Map errors to actions:

  • 429 / rate limit → wait + retry, then fallback
  • Timeout → retry with shorter timeout, then fallback
  • 5xx → retry once, then fallback
  • Invalid output → retry same model with stricter prompt or schema, then fallback
  • Auth / permission errors → do not retry; fail fast and alert

4) Keep request compatibility in mind

Fallback models may differ in:

  • context window
  • tool-calling support
  • structured output reliability
  • temperature/parameter support

So your fallback rule should also specify:

  • whether to truncate context
  • whether to disable tools
  • whether to relax schema validation
  • whether to downgrade to a simpler prompt

5) Add circuit breakers

If a provider is failing repeatedly, stop sending traffic to it for a cooldown period.

Typical behavior:

  • After N failures in M minutes, open the circuit
  • Route all requests to fallbacks
  • Periodically probe the primary provider
  • Close the circuit when healthy again

This avoids hammering a broken provider.

6) Prefer deterministic routing

Don’t make fallback decisions ad hoc in code scattered everywhere. Use a central policy:

  • ordered provider list
  • per-error retry counts
  • cooldown durations
  • request class priorities

For example:

{
  "primary": "openai:gpt-4.1",
  "fallbacks": [
    {"provider": "anthropic:claude-3.5-sonnet", "on": ["timeout", "rate_limit", "server_error"]},
    {"provider": "openai:gpt-4o-mini", "on": ["timeout", "rate_limit", "server_error", "schema_error"]}
  ],
  "retry": {
    "timeout": 1,
    "rate_limit": 2,
    "server_error": 1
  },
  "circuit_breaker": {
    "failure_threshold": 5,
    "window_seconds": 60,
    "cooldown_seconds": 120
  }
}

7) Log and measure fallback usage

Track:

  • fallback rate
  • reason for fallback
  • latency added by retries
  • success rate per provider
  • user impact

This helps you tune the policy and identify flaky providers.

8) Have graceful degradation

If no provider works:

  • return a partial answer
  • ask the user to retry
  • save the request for later processing
  • fall back to a template response
  • use retrieval/cached content if available

9) If you’re using a framework or gateway

Many LLM gateways/proxies support this directly with:

  • provider failover
  • load balancing
  • weighted routing
  • health checks
  • retries and circuit breakers

If you build it yourself, the main logic is:

  1. call provider
  2. classify error
  3. retry if appropriate
  4. move to next provider in policy
  5. fail gracefully

Example pseudo-code

providers = [
    PrimaryProvider(),
    SecondaryProvider(),
    CheapFallbackProvider(),
]

for provider in providers:
    try:
        return provider.generate(prompt, timeout=10)
    except TransientError:
        continue
    except InvalidSchemaError:
        # maybe retry once with stricter prompt here
        continue

raise RuntimeError("All LLM providers failed")

If you want, I can also give you:

  • a production-grade fallback policy design
  • an example in Python/TypeScript
  • a setup for OpenAI + Anthropic + local model failover

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.