Prompt
How do I route AI requests between OpenAI, Anthropic, and Gemini based on cost and latency?
Latest observation
A practical way to route requests across OpenAI, Anthropic, and Gemini is to use a policy-based router that picks the model per request using:
- Task fit
- Estimated cost
- Observed latency
- Current provider health / availability
- Quality constraints like context length, tool use, or vision
1) Decide what you’re optimizing for
Usually you want a weighted objective like:
- Lowest cost for bulk/simple requests
- Lowest latency for interactive UI
- Highest quality for hard reasoning or long-context tasks
- Best reliability as a fallback
A simple scoring formula can be:
[ score(model) = w_c \cdot cost + w_l \cdot latency + w_q \cdot quality_penalty + w_r \cdot risk ]
Pick the model with the lowest score among those that satisfy hard constraints.
2) Use hard rules first
Before scoring, eliminate models that can’t satisfy the request:
- Context window too small
- Missing required feature (tool use, JSON mode, vision, etc.)
- Provider outage / degraded
- Safety / policy restrictions
- Unsupported region or compliance requirement
Example:
- If prompt > 100k tokens, route away from models that can’t fit it.
- If you need structured output, route only to models with reliable JSON/function calling.
3) Estimate cost per request
For each provider/model, estimate:
[ cost = (input_tokens \times input_price) + (output_tokens \times output_price) ]
If output length is unknown, predict it using:
- historical averages by request type
- user-selected “short/normal/detailed”
- a lightweight predictor model
Also include non-token costs if relevant:
- tool calls
- retrieval
- retries
- image/video processing
4) Estimate latency
Latency should be based on measured p50/p95, not just marketing numbers.
Track per model:
- time to first token
- total completion time
- queue time
- retry rate
- timeout rate
Then estimate expected latency for the current request:
[ latency \approx network + queue + model_compute + output_length_factor ]
For interactive apps, route on:
- time to first token
- p95 total latency
5) A good routing strategy
Option A: Rules + fallback
Good starting point.
Example policy:
- If request is “cheap/simple” → smallest/cheapest model
- If request is “fast/UI” → model with best p50 latency
- If request is “hard reasoning” → highest-quality model
- If primary fails → fallback to next best model
Option B: Weighted scoring
More flexible.
Example:
- Cost-sensitive batch jobs:
70% cost, 20% latency, 10% quality - Chat UI:
30% cost, 50% latency, 20% quality - Enterprise support:
20% cost, 20% latency, 60% quality
Option C: Bandit / adaptive routing
Best when traffic is high enough to learn.
Use a multi-armed bandit:
- Explore different models a small % of the time
- Exploit the best-performing one
- Optimize by request category
This helps because latency/cost change over time.
6) Segment requests by type
Don’t route all requests the same way. Common buckets:
- Tiny/simple: classification, extraction, short rewrite
- Standard chat
- Long context
- Reasoning-heavy
- Tool-using / agentic
- Vision / multimodal
- Batch/offline
Then define a preferred provider set for each bucket.
Example:
- Extraction → cheapest model that supports structured output
- Long-context summarization → model with large context window and stable latency
- Reasoning → model with best accuracy
- UI chat → fastest acceptable model
7) Add fallback and hedging
To reduce tail latency and failures:
- Fallback: if provider A errors or times out, retry on B
- Hedged requests: send to a second provider after a short delay if the first is slow
Example:
- Start with Anthropic for quality
- If no first token in 1.5s, launch Gemini or OpenAI
- Cancel the loser when one responds
Use hedging carefully because it increases cost.
8) Maintain a live capability table
Keep a config table per model/provider with:
- input/output pricing
- context window
- p50/p95 latency
- error rate
- supported features
- current health score
Update continuously from telemetry.
Example shape:
{
"openai:gpt-4.1-mini": {
"input_price": 0.00015,
"output_price": 0.0006,
"context": 128000,
"p50_latency_ms": 900,
"p95_latency_ms": 2500,
"supports_tools": true
},
"anthropic:claude": {
"input_price": 0.00025,
"output_price": 0.00125,
"context": 200000,
"p50_latency_ms": 1100,
"p95_latency_ms": 3000,
"supports_tools": true
},
"gemini:pro": {
"input_price": 0.00010,
"output_price": 0.00040,
"context": 1000000,
"p50_latency_ms": 800,
"p95_latency_ms": 2200,
"supports_tools": true
}
}
9) Example routing logic
Pseudo-code:
def route_request(req, models):
candidates = []
for m in models:
if not m.supports(req.features):
continue
if req.input_tokens > m.context_window:
continue
if m.health < 0.8:
continue
est_cost = req.input_tokens * m.input_price + req.expected_output_tokens * m.output_price
est_latency = m.p50_latency_ms + 0.02 * req.expected_output_tokens
score = (
req.cost_weight * est_cost +
req.latency_weight * est_latency +
req.quality_weight * (1 - m.quality_score)
)
candidates.append((score, m))
if not candidates:
return fallback_model(req)
return min(candidates, key=lambda x: x[0])[1]
10) What to measure in production
Track:
- cost per successful request
- p50/p95 latency
- completion rate
- timeout rate
- retry rate
- user satisfaction / task success
- output quality by request class
Then feed those metrics back into the router.
11) Recommended practical architecture
A solid setup is:
- Gateway service
- Model registry with live prices/capabilities
- Telemetry collector
- Router policy engine
- Fallback/hedging layer
- A/B testing framework
This lets you change routing without changing app code.
12) Simple starting policy
If you want an easy first version:
- Use the cheapest model that satisfies context/feature needs
- If request is marked latency-sensitive, choose the model with best recent p50 latency among qualified models
- If request is marked high-value / hard, choose the best-quality model
- If failure or slow response, fallback to next best
If you want, I can also give you:
- a concrete routing algorithm in Python/TypeScript, or
- a decision table for OpenAI vs Anthropic vs Gemini by use case.