Prompt
What’s the cheapest way to run inference on GPUs at scale?
Latest observation
The cheapest way depends on your workload, but in practice the lowest-cost GPU inference setups usually come from combining these four levers:
- Use the smallest/cheapest GPU that meets latency
- Increase utilization with batching and continuous batching
- Run quantized models
- Use the right deployment option: spot/preemptible or reserved bare-metal/cloud GPUs
Lowest-cost options by situation
1) Best raw $/throughput: Spot/preemptible cloud GPUs
If your serving can tolerate interruptions or has failover:
- AWS Spot, GCP Spot, Azure Spot
- Often 50–90% cheaper than on-demand
- Good for:
- stateless inference
- async jobs
- batch generation
- replicated services with autoscaling and retry
Tradeoff: interruptions, capacity scarcity, more ops work.
2) Best steady-state cost: Reserved/committed GPUs
If you have predictable load:
- 1-year or 3-year committed use discounts
- Often cheaper than on-demand by a lot, though not as cheap as spot
- Good for:
- always-on APIs
- stable traffic
Tradeoff: less flexibility.
3) Cheapest at very high volume: Own GPUs / colocation / dedicated bare metal
If you’re at large scale and utilization is high:
- Buying or leasing GPUs can beat cloud rates
- Especially if you can keep them busy 24/7
- Good for:
- large inference fleets
- latency-sensitive services
- large model serving
Tradeoff: capex, ops, hardware refresh, staffing.
What actually reduces inference cost the most
A. Maximize GPU utilization
This is usually the biggest lever.
Use:
- dynamic batching
- continuous batching
- microbatching
- request coalescing
- KV-cache reuse where applicable
Popular serving stacks:
- vLLM
- TensorRT-LLM
- TGI (Text Generation Inference)
- Triton Inference Server
If your GPU sits idle half the time, your cost per token/request is much worse than it needs to be.
B. Quantize models
Quantization often cuts cost dramatically:
- FP16/BF16: baseline
- INT8: common sweet spot
- INT4 / 4-bit: best savings when quality allows
Benefits:
- lower memory use
- more model replicas per GPU
- higher throughput
- cheaper GPUs may become viable
For LLM serving, 4-bit or 8-bit weights can be a major cost reducer, especially for smaller latency-tolerant use cases.
C. Choose the smallest GPU that fits the model
A lot of cost comes from overprovisioning.
For inference, often cheaper to use:
- consumer/“prosumer” GPUs where acceptable
- smaller datacenter GPUs instead of top-tier ones
- newer GPUs only if they improve throughput enough to offset price
Rule of thumb:
- If your model fits and latency is okay, an L4/A10-class GPU may be far cheaper than an A100/H100
- H100s can be economical only when you truly need their throughput or memory bandwidth
D. Reduce token/output length
For LLMs, cost is heavily driven by:
- prompt length
- generated tokens
- concurrency
Ways to lower cost:
- shorten system prompts
- trim context
- summarize history
- cap max output tokens
- use retrieval instead of stuffing context
Common cheapest architecture patterns
For LLM APIs
A cost-efficient setup often looks like:
- vLLM
- INT4/INT8 quantized model
- continuous batching
- spot GPUs with multiple replicas
- autoscaling based on queue depth / token rate
For batch inference
Cheapest is often:
- spot GPUs
- large batches
- offline processing
- checkpointing/retry
- possibly multi-GPU jobs to maximize throughput
For low-latency online serving
- reserved GPUs
- continuous batching
- quantized model
- keep replicas hot
- maybe use a slightly larger GPU to avoid tail latency issues if it improves overall throughput
Practical ranking for cheapest GPU inference
In many real-world cases:
- Spot GPU + batching + quantization
- Reserved GPU + batching + quantization
- Bare metal / owned GPUs at high utilization
- On-demand cloud GPUs with poor utilization
Quick recommendation
If you want the cheapest path that’s usually sane:
- Use vLLM or TensorRT-LLM
- Quantize to 8-bit or 4-bit if quality allows
- Run on spot/preemptible L4/A10-class GPUs
- Use continuous batching
- Keep enough replicas for interruption tolerance
- Measure cost per 1K tokens or per request, not just GPU-hour price
Important caveat
The “cheapest” GPU depends on:
- model size
- context length
- batchability
- latency SLA
- whether you’re doing LLMs, vision, or embedding models
A GPU that’s cheap per hour can be expensive per token if it can’t batch well or doesn’t fit the model efficiently.
If you want, I can help you figure out the cheapest concrete setup for your model by comparing:
- model size
- throughput target
- latency target
- cloud vs self-hosted
- budget constraints