Prompt

How do I size GPU memory for a 70B model?

Artificial Intelligence / AI Infrastructure1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

A good way to size GPU memory for a 70B model is to break it into four pieces:

  1. Model weights
  2. KV cache / context memory
  3. Runtime overhead
  4. Training/fine-tuning states if applicable

The exact number depends a lot on whether you’re doing inference, fine-tuning, or training, and on the precision/quantization.


1) Model weights: the starting point

A 70B parameter model has about 70 billion parameters.

Approximate weight memory:

  • FP16/BF16: 70B × 2 bytes ≈ 140 GB
  • INT8: 70B × 1 byte ≈ 70 GB
  • 4-bit quantized: 70B × 0.5 bytes ≈ 35 GB

In practice, add a little overhead for quantization scales, metadata, padding, etc., so real numbers are often a bit higher than the raw math.

Rule of thumb

  • FP16/BF16 weights: plan for ~140–160 GB
  • 8-bit weights: plan for ~75–90 GB
  • 4-bit weights: plan for ~40–50 GB

2) KV cache: grows with context length and batch size

For inference, the KV cache often becomes the next big memory consumer.

It scales with:

  • number of layers
  • hidden size / head count
  • context length
  • batch size
  • whether you’re using multi-query/grouped-query attention

For a 70B-class model, KV cache can be:

  • a few GB for short contexts and batch size 1
  • tens of GB for long contexts or higher batch sizes

Practical rule of thumb

If you want to serve:

  • 1 user, ~4k context: budget ~2–8 GB extra
  • multiple concurrent users / 8k–32k context: budget ~10–40+ GB extra

This can dominate memory if you’re doing long-context serving.


3) Runtime overhead

Frameworks need extra memory for:

  • CUDA workspace
  • activation buffers
  • temporary tensors
  • fragmentation / allocator overhead

A safe buffer is usually:

  • 10–20% of total model+KV memory, or at least
  • a few GB per GPU

4) If training or fine-tuning, memory is much larger

Full training

Training a 70B model in FP16 is enormous because you also need:

  • gradients
  • optimizer states (e.g. Adam has momentum + variance)
  • activation checkpointing buffers

A rough estimate:

  • Weights: 140 GB
  • Gradients: 140 GB
  • Optimizer states: often 2–8× parameter size depending on optimizer/precision
  • Activations: depends on batch/sequence length

So full training can easily require hundreds of GB to multiple TB of GPU memory, usually across many GPUs.

Fine-tuning

  • LoRA / QLoRA: much smaller, often feasible on a few GPUs or even one large GPU
  • Full fine-tuning: still very memory intensive

Quick sizing examples

Inference examples

70B in FP16

  • Weights: ~140 GB
  • KV cache + overhead: ~10–30 GB
  • Total: ~150–180+ GB

So you’d typically need:

  • 2×80 GB GPUs minimum for comfortable inference
  • or more if long context / high concurrency

70B in 4-bit

  • Weights: ~35–50 GB
  • KV cache + overhead: ~10–30 GB
  • Total: ~45–80 GB

This may fit on:

  • 1×80 GB GPU for modest context/concurrency
  • 2×48 GB GPUs depending on serving stack and context length

Simple sizing formula

For inference:

Total GPU memory ≈ weight size + KV cache + overhead

Where:

  • weight size ≈ parameters × bytes per parameter
  • KV cache ≈ context-dependent, often the variable that surprises people
  • overhead ≈ 5–20 GB depending on runtime and fragmentation

Practical recommendations

If you want a safe single-number estimate:

  • FP16 70B inference: plan on 160–180 GB total
  • 4-bit 70B inference: plan on 50–80 GB total

If you’re buying GPUs:

  • 1×80 GB: feasible for 4-bit inference, not FP16
  • 2×80 GB: good for FP16 inference and larger contexts
  • 4×48 GB: also a common way to host large models

Best way to be precise

Tell me:

  1. Inference, fine-tuning, or training?
  2. Precision: FP16/BF16, INT8, 4-bit?
  3. Max context length
  4. Batch size / concurrency
  5. Model architecture if known

And I can estimate the GPU memory much more accurately.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.