Prompt
I don’t want to be stuck if an AI provider cuts off API access. What do people use to route between model APIs?
Latest observation
People use LLM gateways / proxies / routers that sit between your app (or agent) and multiple model providers. You point your code at one unified (usually OpenAI-compatible) endpoint, and the gateway handles routing, automatic failover, load balancing, retries, and cost tracking. If one provider goes down, rate-limits you, or cuts access, traffic shifts to another without changing your application code.
The most commonly recommended options in 2026 are:
- LiteLLM — The default open-source self-hosted choice for most teams. It supports 100+ providers behind a single OpenAI-compatible API, with built-in fallbacks, load balancing across keys/models, spend tracking, and virtual keys. Easy to run as a proxy or library.
- OpenRouter — Popular managed/hosted aggregator. One API key gives access to hundreds of models from many providers, with automatic failover and consolidated billing. Lowest operational overhead if you don’t want to run infrastructure.
- Portkey — Strong on production features (guardrails, PII redaction, observability, routing policies). Available as open-source or managed; good when you need more than basic failover.
- Bifrost — High-performance open-source gateway (Go-based) focused on very low latency, weighted load balancing, health-aware routing, and cross-provider fallback chains. Often praised for production reliability under load.
- Other solid options people use: Helicone (observability + routing), Cloudflare AI Gateway or Vercel AI Gateway (if you’re already in those ecosystems, often with zero or low markup), Apache APISIX / Kong AI Gateway (enterprise API-gateway style control), and smaller tools like Hermes-router for specific agent or local-model use cases.
Typical setup pattern:
Configure primary model + ordered fallbacks (or weighted distribution) in the gateway.
- Keep using the familiar OpenAI (or Anthropic) SDK — just change the base URL and key. Add budget limits, rate limiting, or semantic caching if needed.
Optionally route by cost, latency, task type, or request complexity.
Self-hosting (LiteLLM, Bifrost, Portkey OSS) keeps prompts and keys fully under your control. Hosted options (OpenRouter, Portkey cloud, Cloudflare/Vercel) trade a bit of control for zero ops.
This approach is the standard way teams avoid single-provider lock-in or outages.
Brands
Citations
- https://www.digitalapi.ai/blogs/what-is-an-llm-gateway
- https://www.getmaxim.ai/articles/top-5-platforms-for-load-balancing-and-failover-across-ai-model-apis/
- https://www.getmaxim.ai/articles/best-ai-gateway-for-routing-between-openai-anthropic-and-gemini/
- https://www.datadoghq.com/blog/ai-gateways-best-practices/
- https://techsy.io/en/blog/best-llm-gateway-tools
- https://awesomeagents.ai/tools/best-llm-gateway-routing-tools-2026/
- https://wetheflywheel.com/en/ai-model-access/openrouter-alternatives/
- https://github.com/Shaf2665/Hermes-router
- https://apisix.apache.org/ai-gateway/
- https://opper.ai/blog/best-llm-gateways