Chuyển đến nội dung chính

第 31 課:Kubernetes 調試和故障排除

調試 Kubernetes:kubectl debug、臨時容器、kubectl 事件、kubectl top。使用 Cilium Hubble 排除 Pod 故障、節點問題、網路問題。常見問題及詳細解決。

🔒 DevSecOps — 第 31 課 第 31 課:調試與故障排除 KUBERNETES

Kubernetes:從基礎到高級

Module 7: Observability & Monitoring

xdev.asia

🎯 課程目標

掌握調試 Kubernetes 的技巧:從 Pod 故障、節點問題到網路問題。使用 kubectl debug、臨時容器和 Cilium Hubble 進行診斷。

1. kubectl debug-臨時容器

# Attach ephemeral container vào pod đang chạy
# Hữu ích khi Pod dùng distroless image không có shell
kubectl debug -it nginx-pod \
  --image=busybox:1.36 \
  --target=nginx \          # share process namespace với container nginx
  -- sh

Trong ephemeral container:

ps aux # xem processes của nginx ls /proc/1/root/etc/nginx # xem files của container nginx wget -O- http://localhost # test

Debug Node

kubectl debug node/worker-1 -it --image=ubuntu -- bash

Có thể mount host filesystem

ls /host/etc/kubernetes

Copy pod để debug (tạo pod mới với debug image thay thế)

kubectl debug nginx-pod
--copy-to=nginx-debug
--image=nginx:debug
--share-processes
-it

2. Pod 故障排查

2.1 Pod 卡在 Pending 狀態

kubectl describe pod mypod -n production
# Xem Events section:
# Warning FailedScheduling: 0/3 nodes are available:
# 3 Insufficient cpu.

Nguyên nhân phổ biến:

- Không đủ CPU/Memory → giảm requests hoặc thêm node

- Node selector/affinity không match → kiểm tra labels

- Taint không có toleration → thêm toleration

- PVC không bound → kiểm tra StorageClass, PV available

2.2 影像拉回關閉

kubectl describe pod mypod -n production
# Events:
# Warning Failed: Failed to pull image "myregistry.io/myapp:v1":
# unauthorized: authentication required

Fix: tạo imagePullSecret

kubectl create secret docker-registry registry-creds
--docker-server=myregistry.io
--docker-username=myuser
--docker-password=mypassword
-n production

Thêm vào pod spec:

imagePullSecrets:

- name: registry-creds

2.3 崩潰循環回退

# Xem logs của lần chạy trước (container đã crash)
kubectl logs mypod -n production --previous

Xem logs real-time

kubectl logs mypod -n production -f

Xem logs của container cụ thể trong multi-container pod

kubectl logs mypod -n production -c mycontainer --previous

Nguyên nhân phổ biến:

- Application error → fix code, kiểm tra config

- OOMKilled → tăng memory limit

- Liveness probe fail quá sớm → tăng initialDelaySeconds

- readOnlyRootFilesystem: app cần ghi file → thêm emptyDir volume

2.4 OOM被殺死

kubectl describe pod mypod
# State: Terminated
# Reason: OOMKilled

Kiểm tra memory usage

kubectl top pod mypod -n production

Fix: tăng memory limit

kubectl set resources deployment myapp --limits=memory=512Mi -n production

Hoặc xem JVM heap nếu là Java app:

kubectl exec mypod -- jcmd 1 VM.native_memory | head -20

3. 節點故障處理

# Xem trạng thái nodes
kubectl get nodes
# STATUS: NotReady → node có vấn đề

kubectl describe node worker-1

Conditions:

MemoryPressure: True → node sắp hết memory

DiskPressure: True → node sắp hết disk

PIDPressure: True → quá nhiều processes

Ready: False → kubelet không healthy

Check kubelet logs trên node

ssh worker-1 journalctl -u kubelet -n 100 --no-pager

Check containerd

systemctl status containerd journalctl -u containerd -n 50 --no-pager

Disk usage

df -h du -sh /var/lib/containerd/*

4. 網路調試

# Test DNS resolution
kubectl run -it --rm debug --image=busybox:1.36 --restart=Never \
  -- nslookup kubernetes.default
kubectl run -it --rm debug --image=busybox:1.36 --restart=Never \
  -- nslookup backend-service.production.svc.cluster.local

Test Service connectivity

kubectl run -it --rm debug --image=busybox:1.36 --restart=Never
-- wget -qO- http://backend-service:8080/health

Kiểm tra endpoints

kubectl get endpoints backend-service -n production

Nếu ENDPOINTS là <none> → Pod selector không match Service selector

Xem EndpointSlices

kubectl get endpointslices -n production -l kubernetes.io/service-name=backend-service

Debug với Cilium Hubble

hubble observe --namespace production --verdict DROPPED hubble observe --namespace production --pod backend-pod --since 5m hubble observe --from-pod frontend-pod --to-pod backend-pod

5. kubectl Events-重要的資訊來源

# Xem events sorted theo thời gian
kubectl get events --sort-by='.lastTimestamp' -n production

Xem events của pod cụ thể

kubectl events --for pod/mypod -n production

Watch events real-time

kubectl get events -n production --watch

Xem warning events

kubectl get events -n production --field-selector type=Warning

6. kubectl top-資源使用狀況

# Cần metrics-server
kubectl top nodes
kubectl top pods -n production

Sort by CPU

kubectl top pods -n production --sort-by=cpu

Xem tất cả namespaces

kubectl top pods --all-namespaces

Xem containers trong pod

kubectl top pod mypod -n production --containers

7. 應用程式緩慢 — 效能故障排除

# CPU throttling
# Xem cgroup CPU stats
kubectl exec mypod -n production -- cat /sys/fs/cgroup/cpu/cpu.stat
# throttled_time lớn → container bị throttle nhiều

Tăng CPU limit hoặc giảm CPU request để đặt đúng

Memory analysis

kubectl exec mypod -- cat /sys/fs/cgroup/memory/memory.usage_in_bytes kubectl exec mypod -- cat /sys/fs/cgroup/memory/memory.limit_in_bytes

Network latency

kubectl exec mypod -- ping -c 10 backend-service kubectl exec mypod -- time wget -qO- http://backend-service:8080/api

Xem connection tracking

kubectl exec mypod -- cat /proc/net/nf_conntrack | wc -l

8. 常見問題清單

Symptom                     First thing to check
──────────────────────────────────────────────────────────────────
Pod Pending                 kubectl describe pod → Events
Pod CrashLoopBackOff        kubectl logs --previous
Pod ImagePullBackOff        Image name, registry creds
Service not reachable       kubectl get endpoints
DNS not working             kubectl exec -- nslookup
Node NotReady               kubectl describe node → Conditions
                            journalctl -u kubelet on node
Slow requests               kubectl top, CPU throttling, network
PVC not bound               kubectl describe pvc, StorageClass

9. Runbook 範例 — CrashLoopBackOff

#!/bin/bash
# Diagnose CrashLoopBackOff

POD=$1 NS=${2:-default}

echo "=== Pod Status ===" kubectl get pod $POD -n $NS -o wide

echo "=== Pod Events ===" kubectl describe pod $POD -n $NS | grep -A 30 Events

echo "=== Current Logs ===" kubectl logs $POD -n $NS 2>/dev/null || echo "No logs (container not started)"

echo "=== Previous Logs ===" kubectl logs $POD -n $NS --previous 2>/dev/null || echo "No previous logs"

echo "=== Resource Usage ===" kubectl top pod $POD -n $NS 2>/dev/null || echo "metrics-server not available"

echo "=== Node Status ===" NODE=$(kubectl get pod $POD -n $NS -o jsonpath='{.spec.nodeName}') kubectl describe node $NODE | grep -A 10 Conditions

總結

  • kubectl debug:無發行鏡像的臨時容器、調試節點
  • CrashLoopBackOff:參見 kubectl 日誌 --previous
  • 待定:請參閱中的事件 kubectl 描述 pod
  • 網路問題:檢查端點,使用 Hubble 查看丟棄的資料包
  • 效能:kubectl top,檢查 cgroup stats 中的 CPU 限制
  • 活動: kubectl 取得事件 --sort-by='.lastTimestamp' 是最重要的工具