Prompt

What does cost per token look like on different GPU providers?

Artificial Intelligence / AI Infrastructure1 observationLast seen Oct 11, 2026

Latest observation

Oct 11, 2026OpenAI APIWeb search: off

Cost per token on GPU providers depends a lot on:

  • Model size (7B vs 70B vs 405B, etc.)
  • Throughput (tokens/sec per GPU or per instance)
  • Utilization (how busy the GPU is)
  • Batching / context length
  • Whether you’re doing inference or fine-tuning
  • Instance pricing model (on-demand, spot, reserved, serverless)

A useful way to think about it is:

cost per token ≈ hourly GPU cost ÷ tokens generated per hour

or

cost per token ≈ hourly infra cost ÷ throughput

Rough comparison by provider type

1) Hyperscalers

Examples: AWS, GCP, Azure

  • Usually most expensive raw GPU hourly rates
  • Strong enterprise features, availability, networking, compliance
  • Good for production-scale deployments where reliability matters

Typical effect: higher cost/token unless you achieve very high utilization.

2) GPU marketplaces / clouds

Examples: CoreWeave, Lambda, RunPod, Paperspace, Crusoe

  • Often cheaper than hyperscalers
  • Better prices for dedicated inference/training clusters
  • More flexible, sometimes less managed

Typical effect: better cost/token, especially for sustained workloads.

3) Spot / preemptible instances

Available on hyperscalers and some specialized providers

  • Lowest cost
  • But interruptions can hurt throughput and operational simplicity

Typical effect: best raw cost/token if you can tolerate interruptions.

4) Serverless inference providers

Examples: Together, Fireworks, Groq, Modal, Replicate, OpenRouter-backed routing

  • You pay per token or per unit of compute
  • No infrastructure management
  • Sometimes higher markup than running your own GPU, but easier operationally

Typical effect: can be cost-effective for variable workloads, but not always cheapest at scale.


Very rough token-cost ranges

These are order-of-magnitude estimates, not exact quotes:

Small/efficient models (7B–8B class)

  • On cheap GPUs, well-optimized inference can be around:
    • $0.000001 to $0.00002 per output token
    • i.e. about $1 to $20 per million output tokens
  • Input tokens are usually cheaper than output tokens in efficient batched setups, but not always by much.

Mid-size models (13B–34B)

  • Roughly:
    • $0.00001 to $0.00008 per output token
    • i.e. about $10 to $80 per million output tokens

Large models (70B+)

  • Roughly:
    • $0.00005 to $0.0005+ per output token
    • i.e. about $50 to $500+ per million output tokens

These ranges vary massively depending on:

  • quantization (FP16 vs INT8/4-bit)
  • batching
  • KV cache efficiency
  • token generation speed
  • context length
  • whether the model is hosted on a high-end GPU like H100 vs a more modest A10/L4

How different providers tend to compare

AWS / Azure / GCP

  • Best when:
    • you need compliance, private networking, enterprise support
    • you’re already committed to the cloud
  • Usually:
    • higher cost/token
    • but can be competitive if using spot and high utilization

CoreWeave

  • Often strong for:
    • large GPU clusters
    • inference at scale
    • training
  • Often:
    • better cost/token than hyperscalers
    • especially if you can keep GPUs busy

Lambda

  • Often attractive for:
    • research, training, smaller inference deployments
  • Usually:
    • cheaper than hyperscalers
    • good price/performance for NVIDIA GPUs

RunPod / Vast / similar marketplaces

  • Often:
    • very low raw cost
    • good for experiments, bursty workloads
  • Tradeoffs:
    • more variability in reliability/performance
    • sometimes more ops burden

Groq / specialized inference hardware

  • Can be excellent for:
    • ultra-low-latency inference
  • Their cost/token can be very strong for supported models, but comparison isn’t apples-to-apples with general-purpose GPUs.

What matters more than provider name

If you want the cheapest cost/token, the biggest levers are:

  1. Pick a smaller model
  2. Use quantization
  3. Increase batching
  4. Keep GPUs highly utilized
  5. Use spot/preemptible where possible
  6. Choose the right GPU for the model
    • L4/A10 for smaller models
    • A100/H100 for larger or high-throughput workloads
  7. Optimize context length
    • long prompts increase cost substantially

Simple rule of thumb

If you’re comparing providers for inference:

  • Cheapest raw cost/token: spot instances on GPU marketplaces
  • Best managed cost/token: specialized inference providers
  • Highest enterprise cost/token: hyperscalers, but with operational benefits

If you want a real estimate

I can help you calculate estimated cost per million tokens for specific setups, for example:

  • AWS g5.xlarge vs RunPod A100
  • L4 vs A10 vs H100
  • 7B vs 70B model
  • input/output token split
  • target requests/sec

If you tell me:

  1. the model,
  2. which GPUs/providers,
  3. expected tokens/sec or latency target,

I can produce a side-by-side cost/token table.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.