Prompt
I need GPU scheduling that avoids stranded capacity and queue buildup
Latest observation
To avoid stranded GPU capacity and queue buildup, you want a scheduler that combines:
- Packing/fragmentation awareness
- Queue fairness and priority
- Elastic scaling / admission control
- Bin-packing with topology awareness
Here’s a practical approach:
What causes the problem
- Stranded capacity: Jobs reserve GPUs in a way that leaves unusable leftovers, e.g. many 1-GPU jobs blocking a 4-GPU node from fitting a 4-GPU job later.
- Queue buildup: Large jobs wait behind lots of smaller jobs, or resources are reserved inefficiently so the queue grows even when GPUs are technically available.
Scheduling strategy to use
1) Use a packing-first scheduler
Prefer placing jobs to minimize fragmentation:
- Pack smaller jobs onto already partially used nodes.
- Avoid spreading a job across nodes unless required.
- Prefer allocating jobs onto nodes where the leftover GPU count matches likely future demand.
This is basically a best-fit / first-fit decreasing style policy for GPUs.
2) Add queue-aware priority
Prevent small jobs from starving large jobs:
- Use age-based priority: jobs gain priority as they wait.
- Use size-aware priority: large jobs get boosted if they’ve been waiting too long.
- Optional: fair-share across users/teams/queues.
A good rule is:
- Schedule by priority first,
- Then choose the placement that causes the least fragmentation.
3) Use backfilling
If a large job is blocked, allow smaller jobs to run only if they don’t delay the large job’s predicted start time. This keeps GPUs busy without increasing queue delays for large jobs.
4) Use gang scheduling for distributed jobs
For multi-GPU jobs that need all GPUs at once:
- Either start them atomically or not at all.
- Do not partially reserve GPUs for too long.
5) Support preemption for lower-priority jobs
If the system has mixed workloads:
- Let low-priority or interruptible jobs be preempted.
- This reduces queue buildup for urgent or large jobs.
6) Make scheduling topology-aware
If GPUs are connected by NVLink / same host / same rack:
- Prefer placing jobs within a node or within a fast interconnect domain.
- Avoid splitting jobs across slow boundaries unless necessary.
A good practical policy
A robust default policy is:
Priority = waiting time + fair-share bonus + job importance
Then among eligible placements:
- choose the placement that
- minimizes leftover fragmentation,
- keeps multi-GPU jobs contiguous,
- preserves future fit for common job sizes.
If you’re using Kubernetes
You’ll likely need:
- A GPU-aware scheduler/extender
- Queueing system such as:
- Kueue
- Volcano
- YuniKorn
- Optional autoscaling:
- Cluster Autoscaler
- Node autoscaling with GPU node pools
For Kubernetes specifically, look for support for:
- gang scheduling
- queueing with quotas
- priority and preemption
- bin-packing placement strategy
If you’re building your own scheduler
Use this loop:
- Maintain job queues by priority class.
- For each scheduling cycle:
- identify jobs that can fit now,
- rank them by urgency/age,
- place them using best-fit packing.
- If a large job is blocked:
- allow backfilling with short jobs only if they won’t delay it.
- Rebalance with preemption if necessary.
Metrics to watch
To know if it’s working, track:
- GPU utilization %
- average queue wait time
- p95 queue wait time
- fragmentation rate
- stranded GPU-hours
- job slowdown
- preemption rate
- number of unschedulable jobs despite free GPUs
Recommended default combination
If you want one simple recommendation:
Use age-based priority + best-fit packing + backfilling + preemption for low-priority jobs.
That usually gives the best balance between:
- high GPU utilization,
- low stranded capacity,
- and shorter queues.
If you want, I can also give you:
- a scheduler design for Kubernetes
- a pseudocode algorithm
- or a policy comparison table for FIFO vs fair-share vs bin-packing vs backfilling.