Prompt
What should I use to host open-source LLMs cheaply?
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
If your goal is cheap hosting for open-source LLMs, the best choice depends on whether you want:
- lowest cost for occasional use
- steady low-latency inference
- easy setup
- GPU access for bigger models
Good cheap options
1. RunPod
- Usually one of the best price/performance options for LLM hosting.
- Good for renting GPU instances by the hour.
- Easy to deploy popular open-source models.
- Great if you want flexible cost and don’t need always-on production infrastructure.
Best for: experimentation, APIs, small-to-medium production workloads.
2. Vast.ai
- Often the cheapest raw GPU market.
- Very cost-effective if you don’t mind a little more setup and variability.
- Prices can be excellent for consumer GPUs like 3090/4090 or A10.
Best for: very low-budget hosting, self-managed deployments.
3. AWS/GCP/Azure spot instances
- Can be cheap if you use spot/preemptible VMs.
- More reliable tooling than random GPU marketplaces, but usually more expensive than Vast/RunPod.
- Good if you already use cloud infrastructure.
Best for: teams needing cloud integration and some cost savings.
4. Hugging Face Inference Endpoints / Spaces
- Easiest to use, but not usually the cheapest.
- Great for demos, prototypes, and simple deployments.
- Spaces are good for lightweight apps; Inference Endpoints for more serious serving.
Best for: convenience over cost.
5. Self-host on a local GPU box
- If you already own a GPU, this is the cheapest long-term option.
- A single used RTX 3090 or 4090 can serve a lot of models for a fraction of cloud cost.
- Best if you have stable demand.
Best for: ongoing usage, maximum cost control.
Model-serving stack to use
For open-source LLMs, the cheapest/easiest serving software is usually:
- vLLM — excellent throughput and popular for API serving
- Ollama — easiest to get started, good for local/smaller deployments
- TGI (Text Generation Inference) — solid production option
- llama.cpp — best for CPU or quantized small models, low memory usage
Practical recommendation
If you want the simplest “cheap but decent” setup:
- Use RunPod + vLLM
- Rent a modest GPU like an A10, L4, or 3090/4090
- Use quantized models when possible
- Keep the instance off when idle if your workload is sporadic
If you want absolute cheapest:
- Vast.ai + llama.cpp/vLLM
- Use quantized models and consumer GPUs
If you want easy deployment:
- Hugging Face Inference Endpoints or RunPod Serverless
Rule of thumb
- Tiny models / prototypes: CPU or small GPU, possibly Ollama or llama.cpp
- Serious API traffic: vLLM on RunPod/Vast
- Lowest cost at scale: your own GPU server or carefully managed spot/market instances
If you tell me:
- what model size you want to host,
- your expected traffic,
- and whether you need an API or just interactive use,
I can recommend the cheapest exact setup.