Prompt
What should I use to control cost and latency across LLM calls?
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
To control cost and latency across LLM calls, the most effective thing to use is a combination of:
1) A gateway / router layer
Use an LLM gateway or orchestration layer to:
- route requests to the cheapest model that can handle them
- switch models based on task complexity
- set per-request budgets and fallbacks
- cache repeated prompts/responses
- collect usage metrics
Examples: LiteLLM, OpenRouter, Azure AI Gateway, LangSmith/LangGraph wrappers, custom API gateway
2) Model routing by task
Don’t send every request to the biggest model. Use:
- small/fast models for classification, extraction, summarization, simple Q&A
- larger models only for hard reasoning, code generation, or low-confidence cases
A common pattern:
- try small model first
- if confidence is low or the task is complex, escalate
3) Caching
Cache at multiple levels:
- prompt/response cache for identical or near-identical requests
- retrieval cache for repeated context chunks
- tool/result cache for expensive API/tool calls
This is one of the best ways to cut both cost and latency.
4) Token budgeting
Control:
- max input context size
- max output tokens
- truncation/summarization of long histories
- conversation memory pruning
Less context = lower cost and faster responses.
5) Streaming and timeouts
For latency:
- enable streaming responses so users see output sooner
- use timeouts and fallbacks to avoid hanging requests
- set retry policies carefully to avoid runaway cost
6) Observability + evaluation
Track:
- tokens in/out
- latency per model
- cost per request
- success rate / escalation rate
- cache hit rate
Then optimize routing with real data.
If you want one practical recommendation
Use:
- LiteLLM or an API gateway for centralized control
- task-based model routing
- response caching
- token limits
- fallbacks/escalation
If you tell me your stack (OpenAI/Anthropic/etc., Python/Node, single app vs multi-service), I can suggest a concrete setup.