Prompt
I'm building on Kubernetes and need help with model serving and GPU scheduling
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
Absolutely — I can help with both model serving on Kubernetes and GPU scheduling.
To get you to a useful design quickly, here’s the practical breakdown:
1) Model serving on Kubernetes
Common options:
- KServe: great for standardized ML inference on K8s, supports autoscaling, canary, batching, transformers, and multiple runtimes.
- Seldon Core: strong for complex inference graphs, model pipelines, explainers, and custom deployments.
- BentoML: simpler developer experience, often easy to package and deploy.
- Ray Serve: good if you need dynamic Python-based serving, ensembles, or online model composition.
- NVIDIA Triton Inference Server: excellent for high-throughput GPU inference, supports TensorRT, ONNX, PyTorch, TensorFlow, dynamic batching.
If you want:
- simple REST/gRPC serving → BentoML or KServe
- high-performance GPU inference → Triton
- custom pipelines / graphs → Seldon or Ray Serve
2) GPU scheduling on Kubernetes
Typical approaches:
- NVIDIA Kubernetes Device Plugin: required for GPU visibility in pods.
- Node labeling + node affinity: schedule workloads to GPU nodes.
- Taints/tolerations: keep generic workloads off GPU nodes.
- MIG (Multi-Instance GPU): partition supported GPUs into slices for better utilization.
- Time-slicing: share GPUs across workloads when isolation is less strict.
- Priority classes / quotas: control fairness among teams/services.
Basic GPU pod requirements
Your pod typically needs:
resources.limits.nvidia.com/gpu: 1- node selector or affinity to GPU nodes
- toleration for GPU node taint if present
Example:
apiVersion: v1
kind: Pod
metadata:
name: gpu-inference
spec:
containers:
- name: app
image: my-inference-image:latest
resources:
limits:
nvidia.com/gpu: 1
nodeSelector:
accelerator: nvidia
3) Recommended production pattern
A common production setup:
- GPU nodes dedicated via taints
- NVIDIA device plugin installed
- Triton or KServe for inference
- Horizontal Pod Autoscaler or KEDA for scaling request-driven services
- Cluster autoscaler to add GPU nodes when demand increases
- MIG if you have A100/H100/L40S class hardware and smaller models
4) Key questions to choose the right setup
If you answer these, I can give you a concrete architecture and YAML:
- What model framework are you serving? (PyTorch / TensorFlow / ONNX / LLM / custom Python)
- Do you need GPU inference or just CPU serving?
- What GPU type do you have? (A10, T4, A100, H100, L4, etc.)
- Expected traffic: low QPS, steady, or bursty?
- Do you need autoscaling?
- Are you deploying one model or many models?
- Do you want a managed-style API like KServe, or something more custom?
5) If you want, I can help with any of these next:
- a reference Kubernetes deployment for model serving
- GPU node pool + taint/toleration setup
- KServe or Triton deployment YAML
- MIG configuration
- autoscaling strategy
- debugging GPU scheduling issues
If you want, send me:
- your Kubernetes distro (EKS/GKE/AKS/on-prem),
- GPU type,
- serving framework,
- and whether you want KServe or Triton,
and I’ll draft a concrete setup for you.