🎯 課程目標
掌握調試 Kubernetes 的技巧:從 Pod 故障、節點問題到網路問題。使用 kubectl debug、臨時容器和 Cilium Hubble 進行診斷。
1. kubectl debug-臨時容器
# Attach ephemeral container vào pod đang chạy # Hữu ích khi Pod dùng distroless image không có shell kubectl debug -it nginx-pod \ --image=busybox:1.36 \ --target=nginx \ # share process namespace với container nginx -- shTrong ephemeral container:
ps aux # xem processes của nginx ls /proc/1/root/etc/nginx # xem files của container nginx wget -O- http://localhost # test
Debug Node
kubectl debug node/worker-1 -it --image=ubuntu -- bash
Có thể mount host filesystem
ls /host/etc/kubernetes
Copy pod để debug (tạo pod mới với debug image thay thế)
kubectl debug nginx-pod
--copy-to=nginx-debug
--image=nginx:debug
--share-processes
-it
2. Pod 故障排查
2.1 Pod 卡在 Pending 狀態
kubectl describe pod mypod -n production # Xem Events section: # Warning FailedScheduling: 0/3 nodes are available: # 3 Insufficient cpu.Nguyên nhân phổ biến:
- Không đủ CPU/Memory → giảm requests hoặc thêm node
- Node selector/affinity không match → kiểm tra labels
- Taint không có toleration → thêm toleration
- PVC không bound → kiểm tra StorageClass, PV available
2.2 影像拉回關閉
kubectl describe pod mypod -n production # Events: # Warning Failed: Failed to pull image "myregistry.io/myapp:v1": # unauthorized: authentication requiredFix: tạo imagePullSecret
kubectl create secret docker-registry registry-creds
--docker-server=myregistry.io
--docker-username=myuser
--docker-password=mypassword
-n productionThêm vào pod spec:
imagePullSecrets:
- name: registry-creds
2.3 崩潰循環回退
# Xem logs của lần chạy trước (container đã crash) kubectl logs mypod -n production --previousXem logs real-time
kubectl logs mypod -n production -f
Xem logs của container cụ thể trong multi-container pod
kubectl logs mypod -n production -c mycontainer --previous
Nguyên nhân phổ biến:
- Application error → fix code, kiểm tra config
- OOMKilled → tăng memory limit
- Liveness probe fail quá sớm → tăng initialDelaySeconds
- readOnlyRootFilesystem: app cần ghi file → thêm emptyDir volume
2.4 OOM被殺死
kubectl describe pod mypod # State: Terminated # Reason: OOMKilledKiểm tra memory usage
kubectl top pod mypod -n production
Fix: tăng memory limit
kubectl set resources deployment myapp --limits=memory=512Mi -n production
Hoặc xem JVM heap nếu là Java app:
kubectl exec mypod -- jcmd 1 VM.native_memory | head -20
3. 節點故障處理
# Xem trạng thái nodes kubectl get nodes # STATUS: NotReady → node có vấn đềkubectl describe node worker-1
Conditions:
MemoryPressure: True → node sắp hết memory
DiskPressure: True → node sắp hết disk
PIDPressure: True → quá nhiều processes
Ready: False → kubelet không healthy
Check kubelet logs trên node
ssh worker-1 journalctl -u kubelet -n 100 --no-pager
Check containerd
systemctl status containerd journalctl -u containerd -n 50 --no-pager
Disk usage
df -h du -sh /var/lib/containerd/*
4. 網路調試
# Test DNS resolution kubectl run -it --rm debug --image=busybox:1.36 --restart=Never \ -- nslookup kubernetes.default kubectl run -it --rm debug --image=busybox:1.36 --restart=Never \ -- nslookup backend-service.production.svc.cluster.localTest Service connectivity
kubectl run -it --rm debug --image=busybox:1.36 --restart=Never
-- wget -qO- http://backend-service:8080/healthKiểm tra endpoints
kubectl get endpoints backend-service -n production
Nếu ENDPOINTS là <none> → Pod selector không match Service selector
Xem EndpointSlices
kubectl get endpointslices -n production -l kubernetes.io/service-name=backend-service
Debug với Cilium Hubble
hubble observe --namespace production --verdict DROPPED hubble observe --namespace production --pod backend-pod --since 5m hubble observe --from-pod frontend-pod --to-pod backend-pod
5. kubectl Events-重要的資訊來源
# Xem events sorted theo thời gian kubectl get events --sort-by='.lastTimestamp' -n productionXem events của pod cụ thể
kubectl events --for pod/mypod -n production
Watch events real-time
kubectl get events -n production --watch
Xem warning events
kubectl get events -n production --field-selector type=Warning
6. kubectl top-資源使用狀況
# Cần metrics-server kubectl top nodes kubectl top pods -n productionSort by CPU
kubectl top pods -n production --sort-by=cpu
Xem tất cả namespaces
kubectl top pods --all-namespaces
Xem containers trong pod
kubectl top pod mypod -n production --containers
7. 應用程式緩慢 — 效能故障排除
# CPU throttling # Xem cgroup CPU stats kubectl exec mypod -n production -- cat /sys/fs/cgroup/cpu/cpu.stat # throttled_time lớn → container bị throttle nhiềuTăng CPU limit hoặc giảm CPU request để đặt đúng
Memory analysis
kubectl exec mypod -- cat /sys/fs/cgroup/memory/memory.usage_in_bytes kubectl exec mypod -- cat /sys/fs/cgroup/memory/memory.limit_in_bytes
Network latency
kubectl exec mypod -- ping -c 10 backend-service kubectl exec mypod -- time wget -qO- http://backend-service:8080/api
Xem connection tracking
kubectl exec mypod -- cat /proc/net/nf_conntrack | wc -l
8. 常見問題清單
Symptom First thing to check
──────────────────────────────────────────────────────────────────
Pod Pending kubectl describe pod → Events
Pod CrashLoopBackOff kubectl logs --previous
Pod ImagePullBackOff Image name, registry creds
Service not reachable kubectl get endpoints
DNS not working kubectl exec -- nslookup
Node NotReady kubectl describe node → Conditions
journalctl -u kubelet on node
Slow requests kubectl top, CPU throttling, network
PVC not bound kubectl describe pvc, StorageClass
9. Runbook 範例 — CrashLoopBackOff
#!/bin/bash # Diagnose CrashLoopBackOffPOD=$1 NS=${2:-default}
echo "=== Pod Status ===" kubectl get pod $POD -n $NS -o wide
echo "=== Pod Events ===" kubectl describe pod $POD -n $NS | grep -A 30 Events
echo "=== Current Logs ===" kubectl logs $POD -n $NS 2>/dev/null || echo "No logs (container not started)"
echo "=== Previous Logs ===" kubectl logs $POD -n $NS --previous 2>/dev/null || echo "No previous logs"
echo "=== Resource Usage ===" kubectl top pod $POD -n $NS 2>/dev/null || echo "metrics-server not available"
echo "=== Node Status ===" NODE=$(kubectl get pod $POD -n $NS -o jsonpath='{.spec.nodeName}') kubectl describe node $NODE | grep -A 10 Conditions
總結
- kubectl debug:無發行鏡像的臨時容器、調試節點
- CrashLoopBackOff:參見
kubectl 日誌 --previous - 待定:請參閱中的事件
kubectl 描述 pod - 網路問題:檢查端點,使用 Hubble 查看丟棄的資料包
- 效能:kubectl top,檢查 cgroup stats 中的 CPU 限制
- 活動:
kubectl 取得事件 --sort-by='.lastTimestamp'是最重要的工具