Prompt
How do I route requests between OpenAI, Anthropic, and Gemini automatically?
Latest observation
A good way to do this is to build a model router: one API endpoint in front, and logic that decides which provider to send each request to based on cost, latency, capability, and reliability.
Common routing strategies
1) Rule-based routing
Use fixed rules such as:
- Cheap/simple tasks → Gemini Flash / smaller OpenAI model
- High-quality reasoning → Anthropic Claude / OpenAI flagship model
- Long-context requests → whichever model supports the needed context length best
- Safety-sensitive prompts → your preferred provider with stronger moderation
Example rule:
- If
input_tokens > 100k, route to a long-context model. - If task is
"code"or"analysis", route to the strongest reasoning model. - If latency target is under 500 ms, route to the fastest model.
2) Score-based routing
Assign each provider a score for the request:
- estimated quality
- estimated cost
- estimated latency
- context fit
- current error rate
Then choose the provider with the highest weighted score.
Example:
score = 0.4*quality - 0.3*cost - 0.2*latency - 0.1*error_rate
3) Fallback routing
Send to a primary model first, then fail over:
- OpenAI → Anthropic → Gemini
- or use one provider per task type, with fallback on timeout/error
This is easy and robust, but not as optimal as dynamic routing.
4) Learned routing
Collect historical data on:
- prompt features
- provider choice
- user satisfaction
- token usage
- latency
- completion success
Then train a classifier or bandit model to predict the best provider.
This works well at scale, but starts more simply with rules.
A practical architecture
Components
-
Request classifier
- Detects task type: chat, code, summarization, extraction, translation, reasoning
- Estimates prompt length and needed context
-
Policy engine
- Applies rules or scores to select a provider
-
Provider adapters
- One adapter each for OpenAI, Anthropic, Gemini
- Normalize request/response formats
-
Fallback + retry
- Timeouts
- Provider errors
- Rate limits
- Automatic downgrade to cheaper/faster model
-
Observability
- Log latency, cost, output quality, retries, and failures
Example routing logic
Simple heuristic
def route_request(prompt, metadata):
tokens = estimate_tokens(prompt)
task = classify_task(prompt)
if tokens > 100_000:
return "gemini-long-context"
if task in ["reasoning", "code", "analysis"]:
return "anthropic-sonnet"
if task in ["summarization", "extraction"] and tokens < 20_000:
return "gemini-flash"
return "openai-default"
With fallback
def call_with_fallback(request):
providers = ["openai", "anthropic", "gemini"]
for p in providers:
try:
return send_to_provider(p, request)
except (TimeoutError, RateLimitError):
continue
raise Exception("All providers failed")
Important normalization details
Different providers have different APIs and features, so normalize:
- message format
- system prompt handling
- tool/function calling
- max tokens
- streaming
- safety settings
Also make sure you handle:
- different stop reasons
- response schemas
- structured output support
- multimodal inputs if needed
How to decide “best” automatically
If you want true automatic routing, use a hybrid approach:
Start with rules
- fast to implement
- predictable
- easy to debug
Add telemetry
Track per-request:
- provider
- task type
- prompt length
- latency
- cost
- user rating / success signal
Then optimize
Use one of:
- multi-armed bandits to explore/exploit provider choices
- offline evaluation using logged data
- A/B tests by segment
Recommended simple policy
A solid default policy:
-
Use cheap model for:
- classification
- extraction
- short summarization
- routine chat
-
Use strong reasoning model for:
- coding
- complex reasoning
- planning
- legal/financial-style drafting
-
Use long-context model for:
- large documents
- many chat turns
- RAG with big retrieved chunks
-
Fallback to alternate provider on:
- timeout
- rate limit
- transient errors
If you want to implement this
I can give you:
- a Python router example
- a Node/TypeScript router
- a LangGraph / LangChain routing setup
- or a production-grade architecture with caching, retries, and metrics
If you want, I can also show a concrete implementation that routes among OpenAI, Anthropic, and Gemini with fallbacks.