Prompt
What should I use for cheap AI inference in an app?
Latest observation
Oct 8, 2026
OpenAI APIWeb search: off
If you want cheap AI inference in an app, the best choice depends on whether you need text, images, speech, or embeddings, but in general:
Good low-cost options
1. Use smaller open-source models via a hosted API
Best when you want easy integration and lower cost than frontier models.
Examples:
- OpenAI: small/mini models for text tasks
- Anthropic / Google / Mistral / Together / Fireworks / Groq: often cheaper options for certain workloads
- Open-source models hosted on providers:
- Llama 3.x small variants
- Mistral 7B / Mixtral
- Qwen small models
Why:
- No infra to manage
- Pay per token/request
- Easy to scale
2. Run an open-source model on your own GPU
Best if you have enough traffic to amortize hardware cost.
Common stack:
- vLLM or TGI for serving
- Llama / Mistral / Qwen models
- A single decent GPU can be much cheaper than API calls at moderate/high volume
Why:
- Lowest marginal cost at scale
- Full control
- Can optimize aggressively
Tradeoff:
- Ops complexity
- You need traffic to justify it
3. Use serverless GPU providers
Good middle ground between hosted APIs and self-hosting.
Examples:
- Modal
- Runpod
- Replicate
- Baseten
- Beam
Why:
- Can be cheaper for bursty workloads
- Less ops than self-hosting
- Often good for prototypes and moderate production use
4. Use distilled/small models
If your task is narrow, this is usually the cheapest.
Examples:
- Tiny LLMs for classification, extraction, routing
- Embedding models instead of full LLMs for search
- Smaller vision models for OCR/classification
Why:
- Much lower inference cost
- Faster latency
- Often enough for app features
Rule of thumb
- Prototype / low traffic: hosted API
- Moderate traffic, bursty: serverless GPU
- High traffic: self-host open-source on GPU
- Simple tasks: use the smallest possible model, or embeddings/classifiers instead of an LLM
Ways to cut cost a lot
- Use smaller models first
- Cache repeated prompts/results
- Trim context aggressively
- Use RAG so the model reads less
- Use streaming only when needed
- Batch requests where possible
- Prefer text-only over multimodal if you can
My default recommendation
If you’re building an app and care about cheap inference:
- Start with a small hosted model
- If cost grows, move to serverless GPU
- If usage is steady and high, self-host an open-source model
If you want, I can give you a specific cheapest option by use case:
- chat app
- image generation
- OCR/extraction
- embeddings/search
- speech-to-text
- classification/routing