🎯 MỤC TIÊU BÀI HỌC
- ✅ Systematic troubleshooting methodology
- ✅ Pod troubleshooting (CrashLoopBackOff, Pending, OOM)
- ✅ Networking issues (DNS, service discovery, connectivity)
- ✅ Storage issues (PV/PVC, mount errors)
- ✅ Control plane troubleshooting
- ✅ Essential diagnostic tools
PHẦN 1: TROUBLESHOOTING METHODOLOGY
Systematic Approach:
1. IDENTIFY → What's the symptom? What changed?
2. ISOLATE → Which component? Which layer?
3. DIAGNOSE → Logs, events, metrics, traces
4. FIX → Apply fix (restart, rollback, patch)
5. VERIFY → Confirm fix, check SLO
6. DOCUMENT → Postmortem, update runbook
Troubleshooting Layers:
┌─────────────────────────────────────┐
│ Application (code, config, deps) │
├─────────────────────────────────────┤
│ Container (image, resources, probes)│
├─────────────────────────────────────┤
│ Pod (scheduling, lifecycle, volumes)│
├─────────────────────────────────────┤
│ Service (DNS, routing, load balance)│
├─────────────────────────────────────┤
│ Node (kubelet, OS, hardware) │
├─────────────────────────────────────┤
│ Cluster (API server, etcd, network) │
└─────────────────────────────────────┘
PHẦN 2: POD TROUBLESHOOTING
# Pod status diagnosis:
# ImagePullBackOff:
kubectl describe pod | grep -A5 "Events"
# Fix: check image name, registry credentials, network
# CrashLoopBackOff:
kubectl logs --previous # Previous container logs
kubectl describe pod # Check exit code
# Common: app error, missing config, wrong command
# Pending:
kubectl describe pod | grep -A10 "Events"
# Common causes:
# - Insufficient resources → check node capacity
# - No matching node selector/affinity
# - PVC not bound
kubectl get events --sort-by=.lastTimestamp
# OOMKilled:
kubectl describe pod | grep "OOMKilled"
kubectl top pod
# Fix: increase memory limits, fix memory leak
# Evicted:
kubectl get pods --field-selector=status.phase=Failed
kubectl describe pod | grep -i evict
# Common: node disk pressure, memory pressure
# Debug containers (ephemeral):
kubectl debug pod/ -it --image=busybox:1.36 --target=app-container
# Debug with network tools:
kubectl debug pod/ -it --image=nicolaka/netshoot --target=app-container
# Copy files from pod for analysis:
kubectl cp :/var/log/app.log ./app.log
PHẦN 3: NETWORKING TROUBLESHOOTING
# DNS resolution:
kubectl run dns-test --image=busybox:1.36 --rm -it -- nslookup order-service
kubectl run dns-test --image=busybox:1.36 --rm -it -- nslookup order-service.default.svc.cluster.local
# Check CoreDNS:
kubectl -n kube-system get pods -l k8s-app=kube-dns
kubectl -n kube-system logs -l k8s-app=kube-dns --tail=50
# Service connectivity:
kubectl run nettest --image=nicolaka/netshoot --rm -it -- bash
# Inside pod:
curl -v http://order-service:8080/health
traceroute order-service
nmap -p 8080 order-service
# Check endpoints:
kubectl get endpoints order-service
# Empty endpoints = no pods match service selector
# NetworkPolicy issues:
kubectl get networkpolicies -n
# Test: temporarily delete NetworkPolicy to confirm it's the cause
# Cilium network debugging:
kubectl -n kube-system exec -it ds/cilium -- cilium status
kubectl -n kube-system exec -it ds/cilium -- cilium monitor
kubectl -n kube-system exec -it ds/cilium -- cilium policy get
PHẦN 4: STORAGE TROUBLESHOOTING
# PVC stuck in Pending:
kubectl describe pvc
# Common causes:
# - No StorageClass matching
# - Ceph cluster full
# - Volume not available in zone
# Check Ceph health:
kubectl -n rook-ceph exec -it deploy/rook-ceph-tools -- ceph status
kubectl -n rook-ceph exec -it deploy/rook-ceph-tools -- ceph osd df
kubectl -n rook-ceph exec -it deploy/rook-ceph-tools -- ceph health detail
# Volume mount errors:
kubectl describe pod | grep -A5 "Warning"
# "Unable to attach or mount volumes"
# Fix: check PV node affinity, RBD map conflicts
# Multi-attach errors (RWO volume):
# Only one node can attach RWO volume
# Fix: delete stuck pod on old node, or use RWX (CephFS)
# Expand PVC:
kubectl patch pvc data-pvc -p '{"spec":{"resources":{"requests":{"storage":"20Gi"}}}}'
# StorageClass must have allowVolumeExpansion: true
PHẦN 5: CONTROL PLANE TROUBLESHOOTING
# API server issues:
kubectl get --raw /healthz
kubectl get --raw /readyz
kubectl get componentstatuses # deprecated but still works
# Check control plane pods:
kubectl -n kube-system get pods
kubectl -n kube-system logs kube-apiserver-master-1
kubectl -n kube-system logs kube-controller-manager-master-1
kubectl -n kube-system logs kube-scheduler-master-1
# etcd health:
ETCDCTL_API=3 etcdctl endpoint health \
--endpoints=https://10.0.1.11:2379,https://10.0.1.12:2379,https://10.0.1.13:2379 \
--cacert=/etc/kubernetes/pki/etcd/ca.crt \
--cert=/etc/kubernetes/pki/etcd/server.crt \
--key=/etc/kubernetes/pki/etcd/server.key
# etcd performance:
ETCDCTL_API=3 etcdctl endpoint status --write-out=table \
--endpoints=... --cacert=... --cert=... --key=...
# Node not ready:
kubectl describe node
# Check conditions: MemoryPressure, DiskPressure, PIDPressure
ssh systemctl status kubelet
ssh journalctl -u kubelet --tail=100
PHẦN 6: COMMON ERRORS QUICK REFERENCE
| Error | Common Cause | Quick Fix |
|---|---|---|
| ImagePullBackOff | Wrong image name/tag, no pull secret | Check image, add imagePullSecrets |
| CrashLoopBackOff | App crash, missing env/config | Check logs --previous |
| Pending | Insufficient resources, no node matches | Check events, node capacity |
| OOMKilled | Memory limit exceeded | Increase limit or fix leak |
| CreateContainerConfigError | Missing ConfigMap/Secret | Check referenced resources exist |
| Evicted | Node disk/memory pressure | Clean up node, increase resources |
| Back-off restarting | Readiness probe failing | Check probe config, port, path |
| connection refused | Service not ready, wrong port | Check endpoints, service ports |
| DNS resolution failed | CoreDNS down, wrong service name | Check coredns pods, FQDN |
💡 KEY TAKEAWAYS
- Systematic approach: Identify → Isolate → Diagnose → Fix → Verify
- kubectl describe: First tool for any K8s issue
- Events: kubectl get events --sort-by=.lastTimestamp
- Debug containers: Ephemeral containers for live diagnosis
- Layer-by-layer: App → Container → Pod → Service → Node → Cluster
- Document: Every fix becomes a runbook entry
🎯 BÀI TẬP
Bài tập 1: Troubleshooting Challenge
- Teammate creates 5 broken deployments (wrong image, missing configmap, etc.)
- Diagnose and fix each using only kubectl
- Document diagnosis steps for each
Bài tập 2: Diagnostic Toolkit
- Create a "debug toolbox" pod with netshoot, pg_isready, redis-cli
- Practice DNS, network, storage troubleshooting
- Build personal troubleshooting cheat sheet
📚 BÀI TIẾP THEO
Trong Bài 50: Capstone Project — E-Commerce Platform, chúng ta sẽ build và deploy toàn bộ hệ thống microservices end-to-end.