Prompt
What GPU instance type should I use for a 13B model?
Latest observation
Oct 11, 2026
OpenAI APIWeb search: off
For a 13B parameter model, the right GPU instance depends mostly on precision and whether you’re doing inference or fine-tuning.
Quick rule of thumb
- Inference in FP16/BF16: needs about 26 GB just for weights, so in practice use 40 GB+ VRAM
- Good choices: A100 40GB, A10G 24GB only if using quantization/offloading, L40S 48GB, H100 80GB
- Inference in 8-bit: often fits in ~14–18 GB, so 24 GB VRAM can work
- Good choices: NVIDIA A10G 24GB, L4 24GB, RTX 4090 24GB
- Inference in 4-bit: often fits in ~7–10 GB
- Good choices: T4 16GB, L4 24GB, A10G 24GB
- Fine-tuning / LoRA: usually want 24–48 GB VRAM
- For comfortable training: A100 40GB/80GB, L40S 48GB, H100 80GB
- For LoRA/QLoRA: A10G 24GB or L4 24GB can work
Practical recommendation
If you want a single safe answer:
- Best all-around: A100 40GB
- Cost-effective inference / QLoRA: A10G 24GB
- Best budget inference for smaller batches: L4 24GB
If you tell me:
- inference or fine-tuning,
- FP16 / 8-bit / 4-bit, and
- your cloud provider (AWS, GCP, Azure, etc.),
I can recommend an exact instance type.