Prompt
How do I decide whether spot GPUs are safe for training?
Latest observation
Spot GPUs can be a great cost saver for training, but they’re “safe” only if your training setup can tolerate interruption. A good way to decide is to evaluate workload tolerance, recovery strategy, and economics.
Use spot GPUs when:
- Your training can resume from checkpoints without losing much progress.
- Interruptions are acceptable and you can requeue jobs automatically.
- You have distributed or elastic training that can handle node loss.
- The cost savings are large enough to justify some operational complexity.
- You can afford delayed completion if capacity disappears.
Avoid spot GPUs when:
- You’re doing one-off training runs that can’t easily be resumed.
- A lost run would be very expensive in researcher time or deadline risk.
- Your training is stateful and brittle (e.g., large in-memory state, hard-to-restart pipelines).
- You need predictable completion time more than low cost.
- You’re in a phase where debugging interruptions would waste a lot of engineering time.
Questions to ask before using spot:
-
Can I checkpoint frequently enough?
Ideally every few minutes or at least often enough that rerunning is cheap. -
How long does restart take?
If booting, loading data, and restoring state take 20–30 minutes, frequent interruptions may erase savings. -
What is my preemption rate?
Estimate how often spot instances are reclaimed in your region/provider and for your GPU type. -
How expensive is failure?
Consider lost compute + engineer time + schedule impact. -
Can I automate recovery?
Good spot setups have:- automatic checkpointing
- job retry logic
- instance termination notices handled gracefully
- artifact/version tracking
A simple rule of thumb
Spot is usually worth it if:
Expected savings > expected interruption cost
A rough way to think about interruption cost:
cost of lost workrestart overheadextra engineering timedeadline risk
If interruptions would cause you to lose only a small fraction of work and restart automatically, spot is often a win.
Practical safeguards
If you do use spot GPUs:
- Save checkpoints to durable storage frequently.
- Make checkpoints small and fast to write.
- Test restoring from a checkpoint before large runs.
- Use job orchestration that retries on failure.
- Track experiment state externally, not only on the instance.
- Prefer training code that can recover from mid-epoch interruptions.
Best use cases
Spot GPUs are especially good for:
- hyperparameter sweeps
- large-scale experimentation
- pretraining where runs are resumable
- batch jobs with many independent tasks
- training pipelines with robust checkpointing
Worst use cases
- interactive debugging
- small but critical production training
- jobs with long warmup and no checkpoints
- experiments that must finish by a fixed time with little slack
If you want, I can help you make a spot-vs-on-demand decision checklist or estimate the break-even point for your specific training job.