Prompt
Do I need H100s for fine-tuning a 13B model?
Latest observation
Oct 11, 2026
OpenAI APIWeb search: off
No — you generally do not need H100s to fine-tune a 13B model.
What you actually need depends on how you fine-tune:
1) Full fine-tuning
If you update all weights, a 13B model is expensive in memory.
- Training in bf16/fp16: model weights alone are ~26 GB
- Add gradients + optimizer states + activations → memory balloons a lot
Practical hardware:
- Usually multiple A100 80GBs or H100s
- Sometimes fewer GPUs if using heavy sharding (FSDP/DeepSpeed ZeRO-3), but still substantial
2) LoRA / QLoRA / PEFT
If you’re doing parameter-efficient fine-tuning, you do not need H100s.
Typical setups:
- 1× A100 40GB/80GB
- 1× RTX 4090 24GB often works for QLoRA with careful batch/seq settings
- Even smaller GPUs can work for some configurations
Rough rule of thumb
- Need H100s? Only if you want fast, large-scale, full fine-tuning or very long context / large batch sizes.
- Don’t need H100s? If you’re doing LoRA/QLoRA, almost certainly not.
If you tell me these 3 things, I can estimate exact GPU needs:
- Full fine-tune vs LoRA/QLoRA
- Max sequence length
- Desired batch size / number of training examples per step