Prompt
How do I get more GPUs when AWS and Azure are sold out?
Latest observation
If AWS and Azure are sold out, the fastest path is usually to broaden where you buy, change what you buy, or rent capacity differently.
1) Try other cloud GPU providers
Look at providers that often have available capacity:
- CoreWeave
- Lambda
- Crusoe
- Nebius
- Paperspace / DigitalOcean GPU offerings
- Vast.ai and RunPod for marketplace-style capacity
- Oracle Cloud sometimes has less competition for certain shapes
- Google Cloud if you haven’t checked it yet
2) Reserve or commit for future capacity
If you need GPUs regularly, ask providers about:
- reserved instances / capacity reservations
- enterprise agreements
- spot + fallback on-demand
- regional flexibility to unlock inventory
A sales rep can sometimes place you on a waitlist or find hidden capacity in another region.
3) Use marketplace or broker channels
For burst demand, marketplace platforms can be quicker than large cloud procurement:
- Vast.ai
- RunPod
- Lambda
- sometimes Hugging Face partners or specialized inference hosts
Tradeoff: more variability in reliability and security posture.
4) Change the GPU model or region
If you only need “a GPU,” not a specific one:
- switch from H100 → A100 → L40S → A10 → T4
- look in less popular regions
- use multi-region deployment
- ask for partial allocations instead of full nodes
Often the bottleneck is the exact SKU, not GPUs in general.
5) Use fewer GPUs more efficiently
If this is for training/inference:
- enable mixed precision
- use quantization
- try LoRA / QLoRA / PEFT
- use gradient checkpointing
- model parallelism only if needed
- batching and speculative decoding for inference
This can cut GPU need significantly.
6) Consider owned hardware or colocation
If the demand is recurring and large:
- buy servers directly from Dell, Supermicro, Lenovo, HPE
- place them in a colo
- use an HPC provider
- lease hardware from a GPU hosting company
This often beats cloud shortages for steady workloads.
7) Ask vendors directly
For Nvidia-based systems, also contact:
- OEMs
- cloud account teams
- GPU hosters
- value-added resellers
They may have allocation before it appears publicly.
8) Watch for preemptible/spot capacity
If your workload can tolerate interruptions:
- use spot instances
- build checkpointing/retry logic
- spread across multiple providers
Practical quick plan
- Check GCP, CoreWeave, Lambda, RunPod, Vast.ai
- Broaden to region + GPU type
- Use spot/marketplace for immediate needs
- Talk to vendor reps for reserved capacity
- If recurring, evaluate colo/owned hardware
If you want, I can help you with a provider shortlist by workload:
- training
- inference
- fine-tuning
- large multi-node clusters
- cheapest available GPUs