Chuyển đến nội dung chính

Bài 10: Troubleshooting Nodes

Debug node NotReady: kubelet, container runtime, certificates. Node conditions, resource pressure, disk pressure. Systematic troubleshooting approach.

Node Troubleshooting Decision Tree — NotReady debug workflow

1. Node Conditions

kubectl describe node node1 | grep -A20 Conditions

Normal state:
  Type              Status  
  ----              ------  
  MemoryPressure    False   ← OK (True = low memory)
  DiskPressure      False   ← OK (True = low disk)
  PIDPressure       False   ← OK (True = too many processes)
  Ready             True    ← Node is healthy

Problem states:
  Ready             False   → kubelet not working
  Ready             Unknown → Node unreachable (network issue)

2. Troubleshoot NotReady Node

Systematic approach — run in order:

1. Check node status
   kubectl get nodes
   kubectl describe node NODE_NAME | tail -40

2. SSH to node
   ssh node1

3. Check kubelet service
   systemctl status kubelet
   journalctl -u kubelet -n 50 --no-pager

4. Check container runtime
   systemctl status containerd
   crictl ps           # List running containers
   crictl pods         # List pod sandboxes

5. Check certificates (common issue after cluster age)
   ls /var/lib/kubelet/pki/
   openssl x509 -in /var/lib/kubelet/pki/kubelet.crt -noout -dates

6. Restart services if needed
   systemctl restart kubelet
   systemctl restart containerd

Exam tip: Quy trình debug NotReady: kubectl describe node → SSH → systemctl status kubelet → journalctl -u kubelet. Hầu hết lỗi: kubelet stopped, wrong API server address, hoặc certificate expired.

3. Common Node Issues

SymptomNguyên nhânFix
Node NotReadykubelet crashedsystemctl restart kubelet
Node UnknownNetwork partitionCheck node network, firewall
MemoryPressure: TrueMemory thiếuEvict pods, scale node
DiskPressure: TrueDisk đầyClean /var/log, /tmp, unused images
Pods stuck TerminatingNode unreachablekubectl delete pod --force --grace-period=0

4. kubelet Configuration

# kubelet config locations
/var/lib/kubelet/config.yaml     # Main config
/etc/kubernetes/kubelet.conf     # kubeconfig (how kubelet connects to API server)
/var/lib/kubelet/kubeconfig      # Alternative path

# Common kubelet config issues:
# Wrong apiserver address
cat /etc/kubernetes/kubelet.conf | grep server

# Wrong cluster DNS
cat /var/lib/kubelet/config.yaml | grep clusterDNS

# Check kubelet's certificate
cat /var/lib/kubelet/config.yaml | grep client-certificate

5. Node Image & Disk Cleanup

# Check disk usage
df -h
du -sh /var/log/*
du -sh /var/lib/containerd

# Clean unused container images
crictl rmi --prune

# Remove old logs
find /var/log/pods -mtime +7 -delete

# Check PID pressure
ps aux | wc -l

6. Cheat Sheet

TaskCommand
Node health summarykubectl describe node NAME
Kubelet statussystemctl status kubelet
Kubelet logsjournalctl -u kubelet -n 100
Running containers on nodecrictl ps
Force delete stuck podkubectl delete pod NAME --force --grace-period=0

7. Practice Questions

Q1: A node shows "Ready: Unknown" status. Which of the following is most likely causing this?

  • A) The kubelet process crashed on the node
  • B) The node cannot be reached by the control plane (network issue) ✓
  • C) All Pods on the node are OOM-killed
  • D) The node has insufficient CPU resources

Explanation: Ready: Unknown means the API server hasn't received a heartbeat from the kubelet recently. This typically indicates node is unreachable (network partition, node powered off). Ready: False means kubelet is reachable but reports a problem.

Q2: After SSH-ing to a NotReady node, you run "systemctl status kubelet" and see "Active: failed". What should you check next?

  • A) kubectl get pods -n kube-system
  • B) journalctl -u kubelet -n 50 to read the error logs ✓
  • C) Delete and recreate the node
  • D) Run kubeadm reset on the node

Explanation: When kubelet fails, journalctl shows the detailed error: certificate issues, wrong API server URL, missing /var/lib/kubelet/config.yaml, etc. This is always the first diagnostic step after confirming kubelet is down.

Q3: A node is reporting DiskPressure: True. What is the immediate effect on workloads?

  • A) All Pods are immediately deleted
  • B) The node is marked unschedulable and BestEffort/Burstable Pods are evicted ✓
  • C) Only new Pod scheduling is prevented
  • D) The kubelet service stops

Explanation: Under disk pressure, Kubernetes triggers pod eviction starting with BestEffort (no requests/limits), then Burstable. Guaranteed Pods are last to be evicted. The node is also tainted to prevent new scheduling.