🎯 Mục tiêu bài học
Nắm vững kỹ năng debugging Kubernetes: từ Pod failures, Node issues đến network problems. Dùng kubectl debug, ephemeral containers, và Cilium Hubble để diagnose.
1. kubectl debug — Ephemeral Containers
# Attach ephemeral container vào pod đang chạy # Hữu ích khi Pod dùng distroless image không có shell kubectl debug -it nginx-pod \ --image=busybox:1.36 \ --target=nginx \ # share process namespace với container nginx -- shTrong ephemeral container:
ps aux # xem processes của nginx ls /proc/1/root/etc/nginx # xem files của container nginx wget -O- http://localhost # test
Debug Node
kubectl debug node/worker-1 -it --image=ubuntu -- bash
Có thể mount host filesystem
ls /host/etc/kubernetes
Copy pod để debug (tạo pod mới với debug image thay thế)
kubectl debug nginx-pod
--copy-to=nginx-debug
--image=nginx:debug
--share-processes
-it
2. Troubleshoot Pod Failures
2.1 Pod Stuck ở Pending
kubectl describe pod mypod -n production # Xem Events section: # Warning FailedScheduling: 0/3 nodes are available: # 3 Insufficient cpu.Nguyên nhân phổ biến:
- Không đủ CPU/Memory → giảm requests hoặc thêm node
- Node selector/affinity không match → kiểm tra labels
- Taint không có toleration → thêm toleration
- PVC không bound → kiểm tra StorageClass, PV available
2.2 ImagePullBackOff
kubectl describe pod mypod -n production # Events: # Warning Failed: Failed to pull image "myregistry.io/myapp:v1": # unauthorized: authentication requiredFix: tạo imagePullSecret
kubectl create secret docker-registry registry-creds
--docker-server=myregistry.io
--docker-username=myuser
--docker-password=mypassword
-n productionThêm vào pod spec:
imagePullSecrets:
- name: registry-creds
2.3 CrashLoopBackOff
# Xem logs của lần chạy trước (container đã crash) kubectl logs mypod -n production --previousXem logs real-time
kubectl logs mypod -n production -f
Xem logs của container cụ thể trong multi-container pod
kubectl logs mypod -n production -c mycontainer --previous
Nguyên nhân phổ biến:
- Application error → fix code, kiểm tra config
- OOMKilled → tăng memory limit
- Liveness probe fail quá sớm → tăng initialDelaySeconds
- readOnlyRootFilesystem: app cần ghi file → thêm emptyDir volume
2.4 OOMKilled
kubectl describe pod mypod # State: Terminated # Reason: OOMKilledKiểm tra memory usage
kubectl top pod mypod -n production
Fix: tăng memory limit
kubectl set resources deployment myapp --limits=memory=512Mi -n production
Hoặc xem JVM heap nếu là Java app:
kubectl exec mypod -- jcmd 1 VM.native_memory | head -20
3. Node Troubleshooting
# Xem trạng thái nodes kubectl get nodes # STATUS: NotReady → node có vấn đềkubectl describe node worker-1
Conditions:
MemoryPressure: True → node sắp hết memory
DiskPressure: True → node sắp hết disk
PIDPressure: True → quá nhiều processes
Ready: False → kubelet không healthy
Check kubelet logs trên node
ssh worker-1 journalctl -u kubelet -n 100 --no-pager
Check containerd
systemctl status containerd journalctl -u containerd -n 50 --no-pager
Disk usage
df -h du -sh /var/lib/containerd/*
4. Network Debugging
# Test DNS resolution kubectl run -it --rm debug --image=busybox:1.36 --restart=Never \ -- nslookup kubernetes.default kubectl run -it --rm debug --image=busybox:1.36 --restart=Never \ -- nslookup backend-service.production.svc.cluster.localTest Service connectivity
kubectl run -it --rm debug --image=busybox:1.36 --restart=Never
-- wget -qO- http://backend-service:8080/healthKiểm tra endpoints
kubectl get endpoints backend-service -n production
Nếu ENDPOINTS là <none> → Pod selector không match Service selector
Xem EndpointSlices
kubectl get endpointslices -n production -l kubernetes.io/service-name=backend-service
Debug với Cilium Hubble
hubble observe --namespace production --verdict DROPPED hubble observe --namespace production --pod backend-pod --since 5m hubble observe --from-pod frontend-pod --to-pod backend-pod
5. kubectl Events — Nguồn thông tin quan trọng
# Xem events sorted theo thời gian kubectl get events --sort-by='.lastTimestamp' -n productionXem events của pod cụ thể
kubectl events --for pod/mypod -n production
Watch events real-time
kubectl get events -n production --watch
Xem warning events
kubectl get events -n production --field-selector type=Warning
6. kubectl top — Resource Usage
# Cần metrics-server kubectl top nodes kubectl top pods -n productionSort by CPU
kubectl top pods -n production --sort-by=cpu
Xem tất cả namespaces
kubectl top pods --all-namespaces
Xem containers trong pod
kubectl top pod mypod -n production --containers
7. Slow Application — Performance Troubleshooting
# CPU throttling # Xem cgroup CPU stats kubectl exec mypod -n production -- cat /sys/fs/cgroup/cpu/cpu.stat # throttled_time lớn → container bị throttle nhiềuTăng CPU limit hoặc giảm CPU request để đặt đúng
Memory analysis
kubectl exec mypod -- cat /sys/fs/cgroup/memory/memory.usage_in_bytes kubectl exec mypod -- cat /sys/fs/cgroup/memory/memory.limit_in_bytes
Network latency
kubectl exec mypod -- ping -c 10 backend-service kubectl exec mypod -- time wget -qO- http://backend-service:8080/api
Xem connection tracking
kubectl exec mypod -- cat /proc/net/nf_conntrack | wc -l
8. Common Issues Checklist
Symptom First thing to check
──────────────────────────────────────────────────────────────────
Pod Pending kubectl describe pod → Events
Pod CrashLoopBackOff kubectl logs --previous
Pod ImagePullBackOff Image name, registry creds
Service not reachable kubectl get endpoints
DNS not working kubectl exec -- nslookup
Node NotReady kubectl describe node → Conditions
journalctl -u kubelet on node
Slow requests kubectl top, CPU throttling, network
PVC not bound kubectl describe pvc, StorageClass
9. Runbook Example — CrashLoopBackOff
#!/bin/bash # Diagnose CrashLoopBackOffPOD=$1 NS=${2:-default}
echo "=== Pod Status ===" kubectl get pod $POD -n $NS -o wide
echo "=== Pod Events ===" kubectl describe pod $POD -n $NS | grep -A 30 Events
echo "=== Current Logs ===" kubectl logs $POD -n $NS 2>/dev/null || echo "No logs (container not started)"
echo "=== Previous Logs ===" kubectl logs $POD -n $NS --previous 2>/dev/null || echo "No previous logs"
echo "=== Resource Usage ===" kubectl top pod $POD -n $NS 2>/dev/null || echo "metrics-server not available"
echo "=== Node Status ===" NODE=$(kubectl get pod $POD -n $NS -o jsonpath='{.spec.nodeName}') kubectl describe node $NODE | grep -A 10 Conditions
Tóm tắt
- kubectl debug: ephemeral containers cho distroless images, debug node
- CrashLoopBackOff: xem
kubectl logs --previous - Pending: xem Events trong
kubectl describe pod - Network issues: kiểm tra endpoints, dùng Hubble để xem dropped packets
- Performance: kubectl top, check CPU throttling trong cgroup stats
- Events:
kubectl get events --sort-by='.lastTimestamp'là công cụ quan trọng nhất