Prompt
How do I monitor cluster health and resource usage in Kubernetes?
Latest observation
To monitor cluster health and resource usage in Kubernetes, you usually combine:
- Built-in Kubernetes tools for quick checks
- Metrics collection for CPU/memory/resource trends
- Logging and alerting for deeper visibility
- Dashboards for ongoing monitoring
1. Basic health checks with kubectl
Check node health
kubectl get nodes
Look for nodes in Ready status.
More detail:
kubectl describe node <node-name>
This shows:
- conditions
- capacity and allocatable resources
- events
- taints
- pod usage info
Check pod health
kubectl get pods -A
Look for:
RunningorCompletedstatus- pods in
CrashLoopBackOff,Pending,Error, orImagePullBackOff
More detail:
kubectl describe pod <pod-name> -n <namespace>
kubectl logs <pod-name> -n <namespace>
Check cluster components
If you’re using a self-managed cluster:
kubectl get componentstatuses
Note: this command is deprecated in newer Kubernetes versions, but still seen in older clusters.
2. Monitor resource usage
Use Metrics Server
To see live CPU/memory usage, install metrics-server if it isn’t already present.
Then use:
kubectl top nodes
kubectl top pods -A
These show:
- node CPU and memory consumption
- pod CPU and memory consumption
You can also inspect a specific namespace:
kubectl top pods -n <namespace>
Compare usage to requests/limits
Check whether workloads are under-provisioned or over-provisioned by looking at:
- resource requests
- resource limits
Example:
kubectl describe pod <pod-name> -n <namespace>
Look for:
Requests:
cpu: 100m
memory: 128Mi
Limits:
cpu: 500m
memory: 256Mi
3. Use monitoring stacks
For production environments, common choices are:
Prometheus + Grafana
- Prometheus collects metrics
- Grafana visualizes them
Typical metrics you can monitor:
- node CPU/memory/disk usage
- pod restarts
- container throttling
- node pressure conditions
- API server latency
- etcd health
- pod scheduling failures
A very common install option is the kube-prometheus-stack Helm chart.
What to alert on
Good alert examples:
- Node not ready
- High CPU or memory usage
- Pod restart rate too high
- Pods stuck pending
- Persistent volume nearing capacity
- API server errors
- etcd unhealthy
4. Check events
Events are often the fastest way to spot issues:
kubectl get events -A --sort-by=.metadata.creationTimestamp
This helps identify:
- scheduling problems
- image pull failures
- OOM kills
- failed mounts
- failed probes
5. Dashboards and UIs
Kubernetes Dashboard
Provides a web UI for:
- cluster overview
- workloads
- namespaces
- resource usage
Grafana dashboards
Best for long-term visibility and alerting.
Useful dashboards include:
- cluster overview
- node exporter dashboard
- namespace/pod resource usage
- deployment health
- API server dashboard
6. Useful commands summary
kubectl get nodes
kubectl get pods -A
kubectl top nodes
kubectl top pods -A
kubectl describe node <node-name>
kubectl describe pod <pod-name> -n <namespace>
kubectl get events -A --sort-by=.metadata.creationTimestamp
7. Best practices
- Set resource requests and limits for all workloads
- Use liveness/readiness probes
- Install metrics-server
- Use Prometheus/Grafana for full observability
- Configure alerts for critical issues
- Review events and logs when troubleshooting
- Track pod restarts and OOMKilled events
If you want, I can also give you:
- a step-by-step setup for Prometheus/Grafana in Kubernetes, or
- a minimal command cheat sheet for checking cluster health quickly.