Prompt
What GPU should I use for batch inference?
Latest observation
Oct 11, 2026
OpenAI APIWeb search: off
It depends on model size, precision, batch size, latency target, and whether you want throughput or cost efficiency. For batch inference, you usually want the cheapest GPU that can keep the model fully loaded and fed at high utilization.
Quick rule of thumb
- Small/medium models (up to ~7B params):
- NVIDIA L4 or RTX 4090 if you’re on a budget and can use consumer hardware.
- Larger models (13B–34B):
- A100 40GB/80GB, L40S, or H100 if you need strong throughput.
- Very large models (70B+):
- Usually A100 80GB, H100 80GB, or multi-GPU with tensor parallelism.
Good choices by goal
Best cost/performance for batch inference
- NVIDIA L4
- Great for inference efficiency
- Good for transformer serving, vision, and smaller LLMs
- Usually more cost-effective than data-center premium GPUs
- NVIDIA L40S
- Strong inference GPU
- Good memory size and bandwidth
- Often excellent for high-throughput batch jobs
Best raw throughput
- H100
- Highest performance
- Best for large models and heavy batching
- Expensive, but fastest
Best value if you can use consumer cards
- RTX 4090
- Very high performance for the price
- Great if your deployment is flexible and you don’t need ECC or datacenter features
- Limited VRAM compared with datacenter GPUs
Best for memory-heavy models
- A100 80GB
- Very useful when model size is the limiting factor
- Mature, reliable choice for large-scale inference
What matters most
- VRAM
- Must fit the model + KV cache + batch data
- If you run out of memory, throughput drops or the model won’t load
- Memory bandwidth
- Often more important than raw TFLOPs for inference
- Tensor core support / FP16 / BF16 / INT8 / FP8
- Quantization can dramatically improve throughput and reduce VRAM use
- Batch size
- Larger batch sizes improve GPU utilization, but increase latency
- Model architecture
- Some models are more memory-bound than compute-bound
Practical recommendations
- If you want one GPU for general batch inference:
- L40S is a very strong default
- If you want lowest cost per token and your models fit:
- L4
- If you want best performance regardless of cost:
- H100
- If you want best value and can tolerate consumer constraints:
- RTX 4090
If you’re running LLMs
Approximate starting points:
- 7B model: 16–24 GB VRAM is usually enough with quantization
- 13B model: 24–40 GB
- 34B model: 48–80 GB
- 70B model: 80 GB+ or multiple GPUs
My default recommendation
If you haven’t narrowed it down yet:
- Choose L40S for a strong balance of throughput, memory, and inference efficiency.
- Choose L4 if cost efficiency is the priority.
- Choose H100 if you need maximum throughput and can afford it.
If you want, I can recommend a GPU based on:
- your model type,
- model size,
- target batch size / throughput,
- and budget.