🎯 Lesson Objective
Understand how to manage node lifecycle in production: maintenance, taints/tolerations, node affinity, topology spread, and automatically detect node problems.
1. Node Lifecycle
# Xem tất cả nodes và status kubectl get nodes -o wide kubectl describe node worker-1Node conditions
kubectl get node worker-1 -o jsonpath='{.status.conditions}' | jq
Ready, MemoryPressure, DiskPressure, PIDPressure, NetworkUnavailable
Xem node resources
kubectl top nodes kubectl describe node worker-1 | grep -A5 "Allocated resources"
2. Cordon and Drain
# Cordon: node ngừng nhận pods mới (pods đang chạy không bị ảnh hưởng) kubectl cordon worker-1 kubectl get node worker-1 # STATUS: Ready,SchedulingDisabledDrain: chuyển pods sang nodes khác (cho maintenance)
kubectl drain worker-1
--ignore-daemonsets \ # bỏ qua DaemonSet pods --delete-emptydir-data \ # xóa emptyDir data --grace-period=60 \ # grace period cho pod termination --timeout=300s # timeout tổngSau maintenance, uncordon để node nhận pods trở lại
kubectl uncordon worker-1
Force drain nếu có pods không evict được
kubectl drain worker-1 --force --ignore-daemonsets --delete-emptydir-data
3. Taints and Tolerations
# Taint nodes kubectl taint nodes gpu-node-1 dedicated=gpu:NoSchedule kubectl taint nodes spot-node-1 cloud.google.com/gke-spot=true:NoSchedule kubectl taint nodes maintenance-node lifecycle=draining:NoExecute # evict ngay cả pods đang chạyRemove taint
kubectl taint nodes gpu-node-1 dedicated=gpu:NoSchedule- # dấu - ở cuối = remove
Xem taints
kubectl describe node gpu-node-1 | grep Taints
# Toleration trong Pod/Deployment
spec:
tolerations:
# Match exact taint
- key: dedicated
operator: Equal
value: gpu
effect: NoSchedule
# Match any taint với key này
- key: cloud.google.com/gke-spot
operator: Exists
effect: NoSchedule
# Tolerate NoExecute với thời gian giới hạn
- key: node.kubernetes.io/not-ready
operator: Exists
effect: NoExecute
tolerationSeconds: 300 # pod ở lại 300s trước khi bị evict
4. Node Affinity
spec:
affinity:
nodeAffinity:
# Bắt buộc: pod chỉ schedule trên nodes này
requiredDuringSchedulingIgnoredDuringExecution:
nodeSelectorTerms:
- matchExpressions:
- key: kubernetes.io/arch
operator: In
values: ["amd64", "arm64"]
- key: node.kubernetes.io/instance-type
operator: In
values: ["c5.2xlarge", "c5.4xlarge"]
# Preferred: ưu tiên schedule trên nodes này (không bắt buộc)
preferredDuringSchedulingIgnoredDuringExecution:
- weight: 100
preference:
matchExpressions:
- key: topology.kubernetes.io/zone
operator: In
values: ["ap-southeast-1a"]
5. Topology Spread Constraints
Distribute pods evenly across zones/nodes to increase availability:
spec:
topologySpreadConstraints:
# Spread đều trên zones
- maxSkew: 1 # chênh lệch tối đa giữa các zones
topologyKey: topology.kubernetes.io/zone
whenUnsatisfiable: DoNotSchedule # hoặc ScheduleAnyway
labelSelector:
matchLabels:
app: my-app
# Spread đều trên nodes
- maxSkew: 2
topologyKey: kubernetes.io/hostname
whenUnsatisfiable: ScheduleAnyway
labelSelector:
matchLabels:
app: my-app
6. Node Problem Detector
# Cài Node Problem Detector (phát hiện node-level issues) kubectl apply -f https://raw.githubusercontent.com/kubernetes/node-problem-detector/master/deployment/node-problem-detector.yamlNPD detect:
- Kernel deadlock (unregister_netdevice hung)
- OOM events
- disk I/O errors
- Container runtime errors
Xem node conditions do NPD thêm vào
kubectl describe node worker-1 | grep -A20 "Conditions"
KernelDeadlock: False
ReadonlyFilesystem: False
FrequentKubeletRestart: False
FrequentDockerRestart: False
NPD tích hợp với cluster autoscaler để thay thế bad nodes tự động
7. Graceful Node Shutdown
# K8s 1.21+: kubelet biết khi node shutdown và gracefully terminate pods # Cấu hình trong /var/lib/kubelet/config.yamlcat <<EOF >> /var/lib/kubelet/config.yaml shutdownGracePeriod: 60s # tổng thời gian cho shutdown shutdownGracePeriodCriticalPods: 10s # thời gian cho critical pods EOF
K8s 1.29+: Non-graceful node shutdown handling
Nếu node crash mà không shutdown gracefully, admin có thể manually trigger taint
kubectl taint node [crashed-node] node.kubernetes.io/out-of-service=nodeshutdown:NoExecute
Pods với PersistentVolumeClaims sẽ bị force-deleted và rescheduled
Sau khi node recover:
kubectl taint node [crashed-node] node.kubernetes.io/out-of-service-
8. Node Auto-Provisioning with Karpenter
# Karpenter (thay Cluster Autoscaler) — chọn instance type phù hợp workload
# NodePool: define node provisioning rules
apiVersion: karpenter.sh/v1
kind: NodePool
metadata:
name: general
spec:
template:
metadata:
labels:
nodepool: general
spec:
nodeClassRef:
apiVersion: karpenter.k8s.aws/v1
kind: EC2NodeClass
name: default
requirements:
- key: karpenter.sh/capacity-type
operator: In
values: ["on-demand", "spot"]
- key: kubernetes.io/arch
operator: In
values: ["amd64", "arm64"]
- key: karpenter.k8s.aws/instance-category
operator: In
values: ["c", "m", "r"] # compute, memory, general
- key: karpenter.k8s.aws/instance-generation
operator: Gt
values: ["3"]
disruption:
consolidationPolicy: WhenEmptyOrUnderutilized
consolidateAfter: 30s # remove empty nodes sau 30s
limits:
cpu: "1000" # max CPU trong pool
memory: 4000Gi
Summary
- Cordon + Drain: standard process for node maintenance
- Taints/Tolerations: controls workload placement (dedicated nodes)
- Node Affinity: binding instance type, arch, zone
- Topology Spread Constraints: ensure pods are evenly distributed across zones/nodes
- Node Problem Detector: detect node-level issues automatically__HTMLTAG_99___
- Karpenter: smarter auto-provisioning than Cluster Autoscaler