Prompt
What should I use to control cost and fallback for LLM calls?
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
Use a routing / orchestration layer with a fallback chain, not direct single-model calls.
What to use
A good setup is:
-
Primary cheap model for most requests
- Fast, low cost
- Handles easy queries
-
Fallback stronger model when needed
- Trigger on low confidence, tool failure, validation failure, or timeout
- Handles complex / ambiguous cases
-
Guardrails / checks
- Output validation
- Confidence scoring
- Budget limits per request/user
- Retry rules with caps
Common patterns
- Model routing: choose model based on request type, length, or risk
- Cascading fallback: cheap model → mid-tier → premium model
- Circuit breaker: stop sending traffic to a failing model temporarily
- Token budgeting: set max input/output tokens
- Response validation: if response fails schema or quality checks, escalate
Tools/frameworks often used
- LangChain / LangGraph
- LlamaIndex
- LiteLLM for model routing and fallbacks
- OpenAI / provider-native routing features if available
- Custom middleware if you want full control
Practical recommendation
If you want something simple and flexible:
- Use LiteLLM or LangGraph
- Configure:
cheap-modelas defaultbetter-modelas fallback- validation + retry limits
- per-request token caps
Example policy
- Use cheap model first
- If:
- output is invalid
- confidence < threshold
- latency > threshold
- tool call fails
- then retry with premium model
If you want, I can sketch a reference architecture or give you a code example for Python/JS.