Chuyển đến nội dung chính

BÀI 49: TROUBLESHOOTING GUIDE

Systematic troubleshooting cho K8s production: pod issues, networking, storage, performance, control plane, common errors, và diagnostic tools.

🔒 DevSecOps — Bài 49 BÀI 49: TROUBLESHOOTING GUIDE

Deploy Microservices On-Premises với Kubernetes HA

Phần 12: Production Operations & Capstone Project

xdev.asia

🎯 MỤC TIÊU BÀI HỌC

  • ✅ Systematic troubleshooting methodology
  • ✅ Pod troubleshooting (CrashLoopBackOff, Pending, OOM)
  • ✅ Networking issues (DNS, service discovery, connectivity)
  • ✅ Storage issues (PV/PVC, mount errors)
  • ✅ Control plane troubleshooting
  • ✅ Essential diagnostic tools

PHẦN 1: TROUBLESHOOTING METHODOLOGY


Systematic Approach:

1. IDENTIFY   → What's the symptom? What changed?
2. ISOLATE    → Which component? Which layer?
3. DIAGNOSE   → Logs, events, metrics, traces
4. FIX        → Apply fix (restart, rollback, patch)
5. VERIFY     → Confirm fix, check SLO
6. DOCUMENT   → Postmortem, update runbook

Troubleshooting Layers:
┌─────────────────────────────────────┐
│ Application (code, config, deps)    │
├─────────────────────────────────────┤
│ Container (image, resources, probes)│
├─────────────────────────────────────┤
│ Pod (scheduling, lifecycle, volumes)│
├─────────────────────────────────────┤
│ Service (DNS, routing, load balance)│
├─────────────────────────────────────┤
│ Node (kubelet, OS, hardware)        │
├─────────────────────────────────────┤
│ Cluster (API server, etcd, network) │
└─────────────────────────────────────┘

PHẦN 2: POD TROUBLESHOOTING

# Pod status diagnosis:

# ImagePullBackOff:
kubectl describe pod  | grep -A5 "Events"
# Fix: check image name, registry credentials, network

# CrashLoopBackOff:
kubectl logs  --previous  # Previous container logs
kubectl describe pod      # Check exit code
# Common: app error, missing config, wrong command

# Pending:
kubectl describe pod  | grep -A10 "Events"
# Common causes:
#   - Insufficient resources → check node capacity
#   - No matching node selector/affinity
#   - PVC not bound
kubectl get events --sort-by=.lastTimestamp

# OOMKilled:
kubectl describe pod  | grep "OOMKilled"
kubectl top pod 
# Fix: increase memory limits, fix memory leak

# Evicted:
kubectl get pods --field-selector=status.phase=Failed
kubectl describe pod  | grep -i evict
# Common: node disk pressure, memory pressure
# Debug containers (ephemeral):
kubectl debug pod/ -it --image=busybox:1.36 --target=app-container

# Debug with network tools:
kubectl debug pod/ -it --image=nicolaka/netshoot --target=app-container

# Copy files from pod for analysis:
kubectl cp :/var/log/app.log ./app.log

PHẦN 3: NETWORKING TROUBLESHOOTING

# DNS resolution:
kubectl run dns-test --image=busybox:1.36 --rm -it -- nslookup order-service
kubectl run dns-test --image=busybox:1.36 --rm -it -- nslookup order-service.default.svc.cluster.local

# Check CoreDNS:
kubectl -n kube-system get pods -l k8s-app=kube-dns
kubectl -n kube-system logs -l k8s-app=kube-dns --tail=50

# Service connectivity:
kubectl run nettest --image=nicolaka/netshoot --rm -it -- bash
# Inside pod:
curl -v http://order-service:8080/health
traceroute order-service
nmap -p 8080 order-service

# Check endpoints:
kubectl get endpoints order-service
# Empty endpoints = no pods match service selector

# NetworkPolicy issues:
kubectl get networkpolicies -n 
# Test: temporarily delete NetworkPolicy to confirm it's the cause

# Cilium network debugging:
kubectl -n kube-system exec -it ds/cilium -- cilium status
kubectl -n kube-system exec -it ds/cilium -- cilium monitor
kubectl -n kube-system exec -it ds/cilium -- cilium policy get

PHẦN 4: STORAGE TROUBLESHOOTING

# PVC stuck in Pending:
kubectl describe pvc 
# Common causes:
#   - No StorageClass matching
#   - Ceph cluster full
#   - Volume not available in zone

# Check Ceph health:
kubectl -n rook-ceph exec -it deploy/rook-ceph-tools -- ceph status
kubectl -n rook-ceph exec -it deploy/rook-ceph-tools -- ceph osd df
kubectl -n rook-ceph exec -it deploy/rook-ceph-tools -- ceph health detail

# Volume mount errors:
kubectl describe pod  | grep -A5 "Warning"
# "Unable to attach or mount volumes"
# Fix: check PV node affinity, RBD map conflicts

# Multi-attach errors (RWO volume):
# Only one node can attach RWO volume
# Fix: delete stuck pod on old node, or use RWX (CephFS)

# Expand PVC:
kubectl patch pvc data-pvc -p '{"spec":{"resources":{"requests":{"storage":"20Gi"}}}}'
# StorageClass must have allowVolumeExpansion: true

PHẦN 5: CONTROL PLANE TROUBLESHOOTING

# API server issues:
kubectl get --raw /healthz
kubectl get --raw /readyz
kubectl get componentstatuses  # deprecated but still works

# Check control plane pods:
kubectl -n kube-system get pods
kubectl -n kube-system logs kube-apiserver-master-1
kubectl -n kube-system logs kube-controller-manager-master-1
kubectl -n kube-system logs kube-scheduler-master-1

# etcd health:
ETCDCTL_API=3 etcdctl endpoint health \
  --endpoints=https://10.0.1.11:2379,https://10.0.1.12:2379,https://10.0.1.13:2379 \
  --cacert=/etc/kubernetes/pki/etcd/ca.crt \
  --cert=/etc/kubernetes/pki/etcd/server.crt \
  --key=/etc/kubernetes/pki/etcd/server.key

# etcd performance:
ETCDCTL_API=3 etcdctl endpoint status --write-out=table \
  --endpoints=... --cacert=... --cert=... --key=...

# Node not ready:
kubectl describe node 
# Check conditions: MemoryPressure, DiskPressure, PIDPressure
ssh  systemctl status kubelet
ssh  journalctl -u kubelet --tail=100

PHẦN 6: COMMON ERRORS QUICK REFERENCE

ErrorCommon CauseQuick Fix
ImagePullBackOffWrong image name/tag, no pull secretCheck image, add imagePullSecrets
CrashLoopBackOffApp crash, missing env/configCheck logs --previous
PendingInsufficient resources, no node matchesCheck events, node capacity
OOMKilledMemory limit exceededIncrease limit or fix leak
CreateContainerConfigErrorMissing ConfigMap/SecretCheck referenced resources exist
EvictedNode disk/memory pressureClean up node, increase resources
Back-off restartingReadiness probe failingCheck probe config, port, path
connection refusedService not ready, wrong portCheck endpoints, service ports
DNS resolution failedCoreDNS down, wrong service nameCheck coredns pods, FQDN

💡 KEY TAKEAWAYS

  1. Systematic approach: Identify → Isolate → Diagnose → Fix → Verify
  2. kubectl describe: First tool for any K8s issue
  3. Events: kubectl get events --sort-by=.lastTimestamp
  4. Debug containers: Ephemeral containers for live diagnosis
  5. Layer-by-layer: App → Container → Pod → Service → Node → Cluster
  6. Document: Every fix becomes a runbook entry

🎯 BÀI TẬP

Bài tập 1: Troubleshooting Challenge

  • Teammate creates 5 broken deployments (wrong image, missing configmap, etc.)
  • Diagnose and fix each using only kubectl
  • Document diagnosis steps for each

Bài tập 2: Diagnostic Toolkit

  • Create a "debug toolbox" pod with netshoot, pg_isready, redis-cli
  • Practice DNS, network, storage troubleshooting
  • Build personal troubleshooting cheat sheet

📚 BÀI TIẾP THEO

Trong Bài 50: Capstone Project — E-Commerce Platform, chúng ta sẽ build và deploy toàn bộ hệ thống microservices end-to-end.