Prompt

How do I route AI requests between OpenAI, Anthropic, and Gemini based on cost and latency?

Artificial Intelligence / AI Infrastructure1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

A practical way to route requests across OpenAI, Anthropic, and Gemini is to use a policy-based router that picks the model per request using:

  1. Task fit
  2. Estimated cost
  3. Observed latency
  4. Current provider health / availability
  5. Quality constraints like context length, tool use, or vision

1) Decide what you’re optimizing for

Usually you want a weighted objective like:

  • Lowest cost for bulk/simple requests
  • Lowest latency for interactive UI
  • Highest quality for hard reasoning or long-context tasks
  • Best reliability as a fallback

A simple scoring formula can be:

[ score(model) = w_c \cdot cost + w_l \cdot latency + w_q \cdot quality_penalty + w_r \cdot risk ]

Pick the model with the lowest score among those that satisfy hard constraints.


2) Use hard rules first

Before scoring, eliminate models that can’t satisfy the request:

  • Context window too small
  • Missing required feature (tool use, JSON mode, vision, etc.)
  • Provider outage / degraded
  • Safety / policy restrictions
  • Unsupported region or compliance requirement

Example:

  • If prompt > 100k tokens, route away from models that can’t fit it.
  • If you need structured output, route only to models with reliable JSON/function calling.

3) Estimate cost per request

For each provider/model, estimate:

[ cost = (input_tokens \times input_price) + (output_tokens \times output_price) ]

If output length is unknown, predict it using:

  • historical averages by request type
  • user-selected “short/normal/detailed”
  • a lightweight predictor model

Also include non-token costs if relevant:

  • tool calls
  • retrieval
  • retries
  • image/video processing

4) Estimate latency

Latency should be based on measured p50/p95, not just marketing numbers.

Track per model:

  • time to first token
  • total completion time
  • queue time
  • retry rate
  • timeout rate

Then estimate expected latency for the current request:

[ latency \approx network + queue + model_compute + output_length_factor ]

For interactive apps, route on:

  • time to first token
  • p95 total latency

5) A good routing strategy

Option A: Rules + fallback

Good starting point.

Example policy:

  • If request is “cheap/simple” → smallest/cheapest model
  • If request is “fast/UI” → model with best p50 latency
  • If request is “hard reasoning” → highest-quality model
  • If primary fails → fallback to next best model

Option B: Weighted scoring

More flexible.

Example:

  • Cost-sensitive batch jobs: 70% cost, 20% latency, 10% quality
  • Chat UI: 30% cost, 50% latency, 20% quality
  • Enterprise support: 20% cost, 20% latency, 60% quality

Option C: Bandit / adaptive routing

Best when traffic is high enough to learn.

Use a multi-armed bandit:

  • Explore different models a small % of the time
  • Exploit the best-performing one
  • Optimize by request category

This helps because latency/cost change over time.


6) Segment requests by type

Don’t route all requests the same way. Common buckets:

  • Tiny/simple: classification, extraction, short rewrite
  • Standard chat
  • Long context
  • Reasoning-heavy
  • Tool-using / agentic
  • Vision / multimodal
  • Batch/offline

Then define a preferred provider set for each bucket.

Example:

  • Extraction → cheapest model that supports structured output
  • Long-context summarization → model with large context window and stable latency
  • Reasoning → model with best accuracy
  • UI chat → fastest acceptable model

7) Add fallback and hedging

To reduce tail latency and failures:

  • Fallback: if provider A errors or times out, retry on B
  • Hedged requests: send to a second provider after a short delay if the first is slow

Example:

  • Start with Anthropic for quality
  • If no first token in 1.5s, launch Gemini or OpenAI
  • Cancel the loser when one responds

Use hedging carefully because it increases cost.


8) Maintain a live capability table

Keep a config table per model/provider with:

  • input/output pricing
  • context window
  • p50/p95 latency
  • error rate
  • supported features
  • current health score

Update continuously from telemetry.

Example shape:

{
  "openai:gpt-4.1-mini": {
    "input_price": 0.00015,
    "output_price": 0.0006,
    "context": 128000,
    "p50_latency_ms": 900,
    "p95_latency_ms": 2500,
    "supports_tools": true
  },
  "anthropic:claude": {
    "input_price": 0.00025,
    "output_price": 0.00125,
    "context": 200000,
    "p50_latency_ms": 1100,
    "p95_latency_ms": 3000,
    "supports_tools": true
  },
  "gemini:pro": {
    "input_price": 0.00010,
    "output_price": 0.00040,
    "context": 1000000,
    "p50_latency_ms": 800,
    "p95_latency_ms": 2200,
    "supports_tools": true
  }
}

9) Example routing logic

Pseudo-code:

def route_request(req, models):
    candidates = []
    for m in models:
        if not m.supports(req.features):
            continue
        if req.input_tokens > m.context_window:
            continue
        if m.health < 0.8:
            continue

        est_cost = req.input_tokens * m.input_price + req.expected_output_tokens * m.output_price
        est_latency = m.p50_latency_ms + 0.02 * req.expected_output_tokens

        score = (
            req.cost_weight * est_cost +
            req.latency_weight * est_latency +
            req.quality_weight * (1 - m.quality_score)
        )
        candidates.append((score, m))

    if not candidates:
        return fallback_model(req)

    return min(candidates, key=lambda x: x[0])[1]

10) What to measure in production

Track:

  • cost per successful request
  • p50/p95 latency
  • completion rate
  • timeout rate
  • retry rate
  • user satisfaction / task success
  • output quality by request class

Then feed those metrics back into the router.


11) Recommended practical architecture

A solid setup is:

  • Gateway service
  • Model registry with live prices/capabilities
  • Telemetry collector
  • Router policy engine
  • Fallback/hedging layer
  • A/B testing framework

This lets you change routing without changing app code.


12) Simple starting policy

If you want an easy first version:

  • Use the cheapest model that satisfies context/feature needs
  • If request is marked latency-sensitive, choose the model with best recent p50 latency among qualified models
  • If request is marked high-value / hard, choose the best-quality model
  • If failure or slow response, fallback to next best

If you want, I can also give you:

  1. a concrete routing algorithm in Python/TypeScript, or
  2. a decision table for OpenAI vs Anthropic vs Gemini by use case.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.