Prompt
I'm building an AI app on Kubernetes — what do I need for serving and observability?
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
For an AI app on Kubernetes, think in two layers:
- Serving: how you expose and scale model inference.
- Observability: how you know it’s fast, correct, reliable, and cost-effective.
1) Serving: what you typically need
Core Kubernetes pieces
- Deployments / StatefulSets
- Deployment for stateless model servers.
- StatefulSet if you need local persistent caches, shards, or tightly coupled replicas.
- Services
- ClusterIP for internal routing.
- LoadBalancer or Ingress for external access.
- Ingress / Gateway API
- For routing, TLS termination, auth integration, and traffic splitting.
- HPA / VPA / KEDA
- HPA for scaling on CPU/GPU/custom metrics.
- KEDA if you want event-driven scaling or queue-based autoscaling.
- Node pools
- Separate CPU and GPU node pools.
- Taints/tolerations and node affinity to keep workloads on the right hardware.
- Resource requests/limits
- Essential for scheduling and avoiding noisy-neighbor issues.
- For GPUs, request extended resources like
nvidia.com/gpu.
Model-serving components
You usually want a dedicated inference layer rather than a generic web app:
- Model server
- Examples: KServe, Seldon, Triton Inference Server, Ray Serve, BentoML, TorchServe, vLLM, TGI.
- Model artifact storage
- Object storage like S3/GCS/Azure Blob or PVC-backed storage.
- Model loading strategy
- Warm startup vs lazy loading.
- Support for model versioning and canary rollout.
- Request batching
- Critical for throughput on GPUs.
- Streaming support
- Useful for LLMs and token-by-token output.
- Caching
- Prompt/prefix caching, embedding cache, response cache where appropriate.
- Queueing/backpressure
- Protects the system under load and helps avoid timeouts and OOMs.
LLM-specific serving concerns
If you’re serving LLMs:
- GPU-aware scheduling
- Bin packing and fragmentation matter a lot.
- KV cache management
- Often the biggest memory pressure point.
- Token streaming
- SSE/WebSocket/gRPC streaming is common.
- Concurrency controls
- Max in-flight requests, max tokens per batch, admission control.
- Quantization
- Can dramatically reduce cost and increase throughput.
- Sharding / tensor parallelism
- Needed for larger models.
- Cold start mitigation
- Pre-pull images, keep warm replicas, preload weights.
Operational features to include
- Health probes
- Liveness, readiness, and startup probes.
- Rollouts
- Blue/green or canary deployments.
- Autoscaling policy
- Scale on queue depth, request latency, GPU utilization, or tokens/sec.
- Timeouts and retries
- Carefully tuned to avoid retry storms.
- Authentication/authorization
- API keys, OAuth, mTLS, or service mesh policies.
- Rate limiting
- Per tenant/user/app to control cost and protect capacity.
2) Observability: what you need to measure
You want observability across three signals:
- Metrics
- Logs
- Traces
And for AI systems, add:
- Model-specific quality/performance metrics
A) Infrastructure metrics
Use Prometheus/Grafana or a managed equivalent.
Track:
- Pod CPU/memory
- GPU utilization
- GPU memory usage
- GPU power/temperature if available
- Network throughput and packet loss
- Disk I/O
- Node pressure / evictions
- Replica counts and restarts
- HPA events and scaling decisions
B) Serving metrics
For inference services, track:
- Request rate
- Latency percentiles: p50, p95, p99
- Error rate
- Timeout rate
- Concurrency / in-flight requests
- Queue depth / queue wait time
- Batch size
- Tokens/sec
- Time to first token for streaming LLMs
- Time spent in prefill vs decode
- Cache hit rates
- Model load time
- OOM / restart counts
C) Model quality metrics
These depend on the application, but common ones are:
- Accuracy / precision / recall / F1
- BLEU / ROUGE / exact match for generation tasks
- Hallucination or groundedness scores
- Human feedback / thumbs up/down
- Abstention rate
- Prompt/response length distributions
- Safety / policy violation rate
- Drift metrics
- Input drift
- Output drift
- Embedding drift
- Retrieval metrics if using RAG:
- Recall@k
- MRR
- Context hit rate
- Citation accuracy
D) Logs
Log:
- Request IDs / correlation IDs
- Model version
- Prompt metadata, not necessarily raw prompt content unless safe/compliant
- Latency breakdown
- Errors and stack traces
- Batch/queue info
- Retry attempts
- Safety filter decisions
- Retrieval source IDs in RAG systems
Be careful with:
- PII
- secrets
- full prompts/completions if sensitive
E) Distributed tracing
Use OpenTelemetry if possible.
Trace the full path:
- API gateway
- auth layer
- retrieval service
- vector DB
- model server
- post-processing
This helps answer:
- Where is latency coming from?
- Is the bottleneck retrieval, serialization, model inference, or downstream IO?
- Which component is failing?
3) Recommended stack
A practical baseline:
Serving
- KServe if you want Kubernetes-native model serving with inference graphs, canary, autoscaling, and model management.
- Triton for optimized GPU inference and multiple framework support.
- vLLM or TGI for LLM serving.
- Ray Serve if you need flexible Python-native serving and orchestration.
Observability
- Prometheus + Grafana for metrics and dashboards
- OpenTelemetry for traces and unified instrumentation
- Loki or ELK/OpenSearch for logs
- Alertmanager for alerting
- Optional managed APM like Datadog/New Relic/Grafana Cloud if you want less ops overhead
AI-specific monitoring
- A model monitoring tool or custom pipeline for:
- input/output capture
- feedback collection
- drift detection
- evaluation jobs
- offline replay testing
4) Architecture tips that save pain
- Separate control plane and data plane
- Keep model management/rollout logic separate from inference traffic.
- Use admission control
- Reject or queue requests when overloaded rather than letting everything time out.
- Precompute where possible
- Embeddings, retrieval indexes, prompt templates.
- Instrument everything with request IDs
- Mandatory for debugging.
- Make model version part of every log/metric
- Essential for rollback and A/B testing.
- Have a fallback path
- Smaller model, cached response, or graceful degradation.
- Plan for GPU scarcity
- Prioritize, queue, or tier requests.
- Security and compliance first
- Secrets management, encryption, audit logs, PII handling.
5) Minimum viable setup
If you want the shortest “good enough” stack:
Serving
- Kubernetes Deployment
- Ingress/Gateway
- Dedicated CPU/GPU node pools
- HPA
- A model server like vLLM/Triton/BentoML
- S3-compatible storage for weights
Observability
- Prometheus
- Grafana dashboards
- OpenTelemetry tracing
- Centralized logs
- Alerts on latency, errors, restarts, GPU memory, queue depth
6) Key questions to decide the right design
If you answer these, the design gets much clearer:
- Are you serving LLMs, vision models, embeddings, or classical ML?
- Do you need GPU inference?
- Is traffic interactive or batch?
- Do you need streaming responses?
- Are you doing RAG?
- What are your SLOs for latency and availability?
- Do you need multi-tenancy and per-tenant quotas?
- How sensitive is the data?
If you want, I can turn this into:
- a reference architecture diagram, or
- a production-ready Kubernetes stack for your specific model type (LLM, RAG, embeddings, vision, etc.).