Prompt
How do I size GPU memory for a 70B model?
Latest observation
A good way to size GPU memory for a 70B model is to break it into four pieces:
- Model weights
- KV cache / context memory
- Runtime overhead
- Training/fine-tuning states if applicable
The exact number depends a lot on whether you’re doing inference, fine-tuning, or training, and on the precision/quantization.
1) Model weights: the starting point
A 70B parameter model has about 70 billion parameters.
Approximate weight memory:
- FP16/BF16: 70B × 2 bytes ≈ 140 GB
- INT8: 70B × 1 byte ≈ 70 GB
- 4-bit quantized: 70B × 0.5 bytes ≈ 35 GB
In practice, add a little overhead for quantization scales, metadata, padding, etc., so real numbers are often a bit higher than the raw math.
Rule of thumb
- FP16/BF16 weights: plan for ~140–160 GB
- 8-bit weights: plan for ~75–90 GB
- 4-bit weights: plan for ~40–50 GB
2) KV cache: grows with context length and batch size
For inference, the KV cache often becomes the next big memory consumer.
It scales with:
- number of layers
- hidden size / head count
- context length
- batch size
- whether you’re using multi-query/grouped-query attention
For a 70B-class model, KV cache can be:
- a few GB for short contexts and batch size 1
- tens of GB for long contexts or higher batch sizes
Practical rule of thumb
If you want to serve:
- 1 user, ~4k context: budget ~2–8 GB extra
- multiple concurrent users / 8k–32k context: budget ~10–40+ GB extra
This can dominate memory if you’re doing long-context serving.
3) Runtime overhead
Frameworks need extra memory for:
- CUDA workspace
- activation buffers
- temporary tensors
- fragmentation / allocator overhead
A safe buffer is usually:
- 10–20% of total model+KV memory, or at least
- a few GB per GPU
4) If training or fine-tuning, memory is much larger
Full training
Training a 70B model in FP16 is enormous because you also need:
- gradients
- optimizer states (e.g. Adam has momentum + variance)
- activation checkpointing buffers
A rough estimate:
- Weights: 140 GB
- Gradients: 140 GB
- Optimizer states: often 2–8× parameter size depending on optimizer/precision
- Activations: depends on batch/sequence length
So full training can easily require hundreds of GB to multiple TB of GPU memory, usually across many GPUs.
Fine-tuning
- LoRA / QLoRA: much smaller, often feasible on a few GPUs or even one large GPU
- Full fine-tuning: still very memory intensive
Quick sizing examples
Inference examples
70B in FP16
- Weights: ~140 GB
- KV cache + overhead: ~10–30 GB
- Total: ~150–180+ GB
So you’d typically need:
- 2×80 GB GPUs minimum for comfortable inference
- or more if long context / high concurrency
70B in 4-bit
- Weights: ~35–50 GB
- KV cache + overhead: ~10–30 GB
- Total: ~45–80 GB
This may fit on:
- 1×80 GB GPU for modest context/concurrency
- 2×48 GB GPUs depending on serving stack and context length
Simple sizing formula
For inference:
Total GPU memory ≈ weight size + KV cache + overhead
Where:
- weight size ≈ parameters × bytes per parameter
- KV cache ≈ context-dependent, often the variable that surprises people
- overhead ≈ 5–20 GB depending on runtime and fragmentation
Practical recommendations
If you want a safe single-number estimate:
- FP16 70B inference: plan on 160–180 GB total
- 4-bit 70B inference: plan on 50–80 GB total
If you’re buying GPUs:
- 1×80 GB: feasible for 4-bit inference, not FP16
- 2×80 GB: good for FP16 inference and larger contexts
- 4×48 GB: also a common way to host large models
Best way to be precise
Tell me:
- Inference, fine-tuning, or training?
- Precision: FP16/BF16, INT8, 4-bit?
- Max context length
- Batch size / concurrency
- Model architecture if known
And I can estimate the GPU memory much more accurately.