Prompt
What does cost per token look like on different GPU providers?
Latest observation
Cost per token on GPU providers depends a lot on:
- Model size (7B vs 70B vs 405B, etc.)
- Throughput (tokens/sec per GPU or per instance)
- Utilization (how busy the GPU is)
- Batching / context length
- Whether you’re doing inference or fine-tuning
- Instance pricing model (on-demand, spot, reserved, serverless)
A useful way to think about it is:
cost per token ≈ hourly GPU cost ÷ tokens generated per hour
or
cost per token ≈ hourly infra cost ÷ throughput
Rough comparison by provider type
1) Hyperscalers
Examples: AWS, GCP, Azure
- Usually most expensive raw GPU hourly rates
- Strong enterprise features, availability, networking, compliance
- Good for production-scale deployments where reliability matters
Typical effect: higher cost/token unless you achieve very high utilization.
2) GPU marketplaces / clouds
Examples: CoreWeave, Lambda, RunPod, Paperspace, Crusoe
- Often cheaper than hyperscalers
- Better prices for dedicated inference/training clusters
- More flexible, sometimes less managed
Typical effect: better cost/token, especially for sustained workloads.
3) Spot / preemptible instances
Available on hyperscalers and some specialized providers
- Lowest cost
- But interruptions can hurt throughput and operational simplicity
Typical effect: best raw cost/token if you can tolerate interruptions.
4) Serverless inference providers
Examples: Together, Fireworks, Groq, Modal, Replicate, OpenRouter-backed routing
- You pay per token or per unit of compute
- No infrastructure management
- Sometimes higher markup than running your own GPU, but easier operationally
Typical effect: can be cost-effective for variable workloads, but not always cheapest at scale.
Very rough token-cost ranges
These are order-of-magnitude estimates, not exact quotes:
Small/efficient models (7B–8B class)
- On cheap GPUs, well-optimized inference can be around:
- $0.000001 to $0.00002 per output token
- i.e. about $1 to $20 per million output tokens
- Input tokens are usually cheaper than output tokens in efficient batched setups, but not always by much.
Mid-size models (13B–34B)
- Roughly:
- $0.00001 to $0.00008 per output token
- i.e. about $10 to $80 per million output tokens
Large models (70B+)
- Roughly:
- $0.00005 to $0.0005+ per output token
- i.e. about $50 to $500+ per million output tokens
These ranges vary massively depending on:
- quantization (FP16 vs INT8/4-bit)
- batching
- KV cache efficiency
- token generation speed
- context length
- whether the model is hosted on a high-end GPU like H100 vs a more modest A10/L4
How different providers tend to compare
AWS / Azure / GCP
- Best when:
- you need compliance, private networking, enterprise support
- you’re already committed to the cloud
- Usually:
- higher cost/token
- but can be competitive if using spot and high utilization
CoreWeave
- Often strong for:
- large GPU clusters
- inference at scale
- training
- Often:
- better cost/token than hyperscalers
- especially if you can keep GPUs busy
Lambda
- Often attractive for:
- research, training, smaller inference deployments
- Usually:
- cheaper than hyperscalers
- good price/performance for NVIDIA GPUs
RunPod / Vast / similar marketplaces
- Often:
- very low raw cost
- good for experiments, bursty workloads
- Tradeoffs:
- more variability in reliability/performance
- sometimes more ops burden
Groq / specialized inference hardware
- Can be excellent for:
- ultra-low-latency inference
- Their cost/token can be very strong for supported models, but comparison isn’t apples-to-apples with general-purpose GPUs.
What matters more than provider name
If you want the cheapest cost/token, the biggest levers are:
- Pick a smaller model
- Use quantization
- Increase batching
- Keep GPUs highly utilized
- Use spot/preemptible where possible
- Choose the right GPU for the model
- L4/A10 for smaller models
- A100/H100 for larger or high-throughput workloads
- Optimize context length
- long prompts increase cost substantially
Simple rule of thumb
If you’re comparing providers for inference:
- Cheapest raw cost/token: spot instances on GPU marketplaces
- Best managed cost/token: specialized inference providers
- Highest enterprise cost/token: hyperscalers, but with operational benefits
If you want a real estimate
I can help you calculate estimated cost per million tokens for specific setups, for example:
- AWS g5.xlarge vs RunPod A100
- L4 vs A10 vs H100
- 7B vs 70B model
- input/output token split
- target requests/sec
If you tell me:
- the model,
- which GPUs/providers,
- expected tokens/sec or latency target,
I can produce a side-by-side cost/token table.