Prompt
I'm unhappy with our current LLM setup because spend is hard to control; what should I use?
Latest observation
Oct 8, 2026
OpenAI APIWeb search: off
If your main pain is uncontrolled spend, the best answer is usually not “pick a cheaper model” but “add a model gateway + routing + budgets + observability.” That gives you much tighter control than using a raw provider directly.
What to use
Best practical setup
Use an LLM gateway/proxy in front of your model providers, such as:
- LiteLLM
- OpenRouter (more for access/routing than internal control)
- Helicone
- Portkey
- AWS Bedrock / Azure OpenAI / Vertex AI if you want stronger enterprise billing controls
These let you:
- set per-team / per-app budgets
- enforce rate limits
- route cheap models first
- fall back to more expensive models only when needed
- log usage by request, user, endpoint, or tag
- monitor token spend in near real time
If you want the strongest spend control
A good pattern is:
- Default to a low-cost model
- e.g. GPT-4o mini, Claude Haiku, Gemini Flash, or a small open model
- Escalate only on failure/low confidence
- use a larger model only for hard cases
- Put hard caps in the gateway
- max tokens, max requests/minute, monthly budget ceilings
- Add caching
- semantic cache for repeated prompts
- Track spend per feature
- so one workflow can’t quietly burn your whole budget
What I’d recommend by scenario
If you want easy control with minimal engineering
- LiteLLM + budget/rate-limit policies
- Add Helicone for observability if needed
If you’re already on a cloud with compliance needs
- AWS Bedrock, Azure OpenAI, or Vertex AI
- These are better for centralized billing and governance
If you want best cost optimization
- Gateway + model routing + caching + evals
- Use smaller models for most requests, bigger ones for exceptions
What not to do
- Don’t let every app call model APIs directly
- Don’t rely on “just monitor invoices”
- Don’t use the biggest model for every request
- Don’t skip token caps
A simple recommendation
If you want a concrete answer:
Use LiteLLM as your gateway, route most traffic to a cheap model, and enforce budgets/rate limits per service.
If you want, I can give you:
- a recommended stack for your exact environment, or
- a decision table comparing LiteLLM vs Bedrock vs OpenRouter vs Helicone.