Prompt
Building a RAG backend and trying to decide what GPU to use
Latest observation
For a RAG backend, the “right” GPU depends mostly on what runs on the GPU:
- Embedding generation
- Reranking
- LLM inference for answer generation
- Batch size / concurrency / latency target
- Model size and context length
If you tell me your model(s), expected QPS, and latency target, I can recommend a specific GPU. In the meantime, here’s a practical way to choose.
Quick recommendation
If you’re mostly doing:
- Embeddings + reranker + small/medium LLMs
- NVIDIA L4: great cost/performance for inference
- NVIDIA A10 / A10G: solid budget option
- NVIDIA L40S: excellent if you want more headroom and stronger throughput
If you want:
- Higher concurrency
- Larger models
- Longer contexts
- Less tuning / fewer memory headaches
- L40S is often the sweet spot
- A100 80GB if you need large models or heavy batching
- H100 if you’re serving large production workloads and cost is secondary
Rule of thumb by workload
1) Embeddings only
If your GPU is mainly for embedding batches:
- L4 is usually enough
- A10G if you need more throughput
- You probably do not need A100/H100
2) Reranking
Rerankers are relatively light:
- L4 / A10G usually fine
- If using a large cross-encoder, a stronger card helps, but still rarely needs top-end GPUs
3) LLM answer generation
This is usually the deciding factor.
Small models
- 7B / 8B models:
- L4 can work
- A10G safer
- L40S much more comfortable
Medium models
- 13B / 14B:
- A10G may work with quantization
- L40S is better
- A100 80GB if you need larger context or throughput
Large models
- 30B+ / 70B:
- Generally A100 80GB or H100
- You may need multi-GPU depending on quantization and context size
Best practical picks
Budget-conscious production RAG
- NVIDIA L4
- Best when:
- using a smaller LLM
- moderate traffic
- embeddings/reranking are important
- Pros: efficient, lower cost
- Cons: limited memory
Balanced production RAG
- NVIDIA L40S
- Best overall choice for many RAG backends
- Pros:
- strong inference performance
- good memory capacity
- easier to scale into larger models and longer contexts
- Cons: costs more than L4/A10G
Heavy production / larger models
- A100 80GB or H100
- Best when:
- serving large models
- high concurrency
- long contexts
- strict latency
- Cons: expensive
Memory matters more than raw FLOPS in RAG
A lot of RAG systems hit memory limits before compute limits because of:
- model weights
- KV cache for long contexts
- batching multiple requests
- reranker + generator on same GPU
If you expect:
- long prompts
- multiple retrieved chunks
- many simultaneous users
then prioritize VRAM heavily.
Simple decision guide
- Prototype / low traffic: L4 or A10G
- Serious production with small-to-mid models: L40S
- Large models or heavy concurrency: A100 80GB
- Maximum throughput and budget isn’t the issue: H100
My default recommendation
If you’re building a general-purpose RAG backend and don’t know your final traffic yet:
Start with an L40S if budget allows. It’s the most “safe” all-around choice for RAG because it handles embeddings, reranking, and a decent range of generation models without immediately running into memory constraints.
If cost is tighter:
Start with an L4 for embeddings/reranking and a smaller LLM, or A10G if you want more breathing room.
Questions that would let me give a precise answer
- Which LLM(s) are you serving? (e.g. Llama 3 8B, 70B, Mistral, etc.)
- Are embeddings generated online or precomputed?
- Do you need reranking?
- Target latency per request?
- Expected concurrent users or QPS?
- Cloud or on-prem? Any budget ceiling?
If you answer those, I’ll recommend a specific GPU and rough architecture.