Chuyển đến nội dung chính

LESSON 49: TROUBLESHOOTING GUIDE

Systematic troubleshooting for K8s production: pod issues, networking, storage, performance, control plane, common errors, and diagnostic tools.

🔒 DevSecOps — Lesson 49 LESSON 49: TROUBLESHOOTING GUIDE

Deploy Microservices On-Premises with Kubernetes HA

Part 12: Production Operations & Capstone Project

xdev.asia

🎯 LESSON OBJECTIVE__HTMLTAG_66___
  • ✅ Systematic troubleshooting methodology
  • ✅ Pod troubleshooting (CrashLoopBackOff, Pending, OOM)
  • ✅ Networking issues (DNS, service discovery, connectivity)
  • ✅ Storage issues (PV/PVC, mount errors)
  • ✅ Control plane troubleshooting__HTMLTAG_77___
  • ✅ Essential diagnostic tools

PART 1: TROUBLESHOOTING METHODOLOGY


Systematic Approach:

1. IDENTIFY   → What's the symptom? What changed?
2. ISOLATE    → Which component? Which layer?
3. DIAGNOSE   → Logs, events, metrics, traces
4. FIX        → Apply fix (restart, rollback, patch)
5. VERIFY     → Confirm fix, check SLO
6. DOCUMENT   → Postmortem, update runbook

Troubleshooting Layers:
┌─────────────────────────────────────┐
│ Application (code, config, deps)    │
├─────────────────────────────────────┤
│ Container (image, resources, probes)│
├─────────────────────────────────────┤
│ Pod (scheduling, lifecycle, volumes)│
├─────────────────────────────────────┤
│ Service (DNS, routing, load balance)│
├─────────────────────────────────────┤
│ Node (kubelet, OS, hardware)        │
├─────────────────────────────────────┤
│ Cluster (API server, etcd, network) │
└─────────────────────────────────────┘

PART 2: POD TROUBLESHOOTING

# Pod status diagnosis:

# ImagePullBackOff:
kubectl describe pod  | grep -A5 "Events"
# Fix: check image name, registry credentials, network

# CrashLoopBackOff:
kubectl logs  --previous  # Previous container logs
kubectl describe pod      # Check exit code
# Common: app error, missing config, wrong command

# Pending:
kubectl describe pod  | grep -A10 "Events"
# Common causes:
#   - Insufficient resources → check node capacity
#   - No matching node selector/affinity
#   - PVC not bound
kubectl get events --sort-by=.lastTimestamp

# OOMKilled:
kubectl describe pod  | grep "OOMKilled"
kubectl top pod 
# Fix: increase memory limits, fix memory leak

# Evicted:
kubectl get pods --field-selector=status.phase=Failed
kubectl describe pod  | grep -i evict
# Common: node disk pressure, memory pressure
# Debug containers (ephemeral):
kubectl debug pod/ -it --image=busybox:1.36 --target=app-container

# Debug with network tools:
kubectl debug pod/ -it --image=nicolaka/netshoot --target=app-container

# Copy files from pod for analysis:
kubectl cp :/var/log/app.log ./app.log

PART 3: NETWORKING TROUBLESHOOTING

# DNS resolution:
kubectl run dns-test --image=busybox:1.36 --rm -it -- nslookup order-service
kubectl run dns-test --image=busybox:1.36 --rm -it -- nslookup order-service.default.svc.cluster.local

# Check CoreDNS:
kubectl -n kube-system get pods -l k8s-app=kube-dns
kubectl -n kube-system logs -l k8s-app=kube-dns --tail=50

# Service connectivity:
kubectl run nettest --image=nicolaka/netshoot --rm -it -- bash
# Inside pod:
curl -v http://order-service:8080/health
traceroute order-service
nmap -p 8080 order-service

# Check endpoints:
kubectl get endpoints order-service
# Empty endpoints = no pods match service selector

# NetworkPolicy issues:
kubectl get networkpolicies -n 
# Test: temporarily delete NetworkPolicy to confirm it's the cause

# Cilium network debugging:
kubectl -n kube-system exec -it ds/cilium -- cilium status
kubectl -n kube-system exec -it ds/cilium -- cilium monitor
kubectl -n kube-system exec -it ds/cilium -- cilium policy get

PART 4: STORAGE TROUBLESHOOTING

# PVC stuck in Pending:
kubectl describe pvc 
# Common causes:
#   - No StorageClass matching
#   - Ceph cluster full
#   - Volume not available in zone

# Check Ceph health:
kubectl -n rook-ceph exec -it deploy/rook-ceph-tools -- ceph status
kubectl -n rook-ceph exec -it deploy/rook-ceph-tools -- ceph osd df
kubectl -n rook-ceph exec -it deploy/rook-ceph-tools -- ceph health detail

# Volume mount errors:
kubectl describe pod  | grep -A5 "Warning"
# "Unable to attach or mount volumes"
# Fix: check PV node affinity, RBD map conflicts

# Multi-attach errors (RWO volume):
# Only one node can attach RWO volume
# Fix: delete stuck pod on old node, or use RWX (CephFS)

# Expand PVC:
kubectl patch pvc data-pvc -p '{"spec":{"resources":{"requests":{"storage":"20Gi"}}}}'
# StorageClass must have allowVolumeExpansion: true

PART 5: CONTROL PLANE TROUBLESHOOTING

# API server issues:
kubectl get --raw /healthz
kubectl get --raw /readyz
kubectl get componentstatuses  # deprecated but still works

# Check control plane pods:
kubectl -n kube-system get pods
kubectl -n kube-system logs kube-apiserver-master-1
kubectl -n kube-system logs kube-controller-manager-master-1
kubectl -n kube-system logs kube-scheduler-master-1

# etcd health:
ETCDCTL_API=3 etcdctl endpoint health \
  --endpoints=https://10.0.1.11:2379,https://10.0.1.12:2379,https://10.0.1.13:2379 \
  --cacert=/etc/kubernetes/pki/etcd/ca.crt \
  --cert=/etc/kubernetes/pki/etcd/server.crt \
  --key=/etc/kubernetes/pki/etcd/server.key

# etcd performance:
ETCDCTL_API=3 etcdctl endpoint status --write-out=table \
  --endpoints=... --cacert=... --cert=... --key=...

# Node not ready:
kubectl describe node 
# Check conditions: MemoryPressure, DiskPressure, PIDPressure
ssh  systemctl status kubelet
ssh  journalctl -u kubelet --tail=100

PART 6: COMMON ERRORS QUICK REFERENCE

ErrorCommon CauseQuick Fix
ImagePullBackOffWrong image name/tag, no pull secret_Check image, add imagePullSecrets
CrashLoopBackOffApp crash, missing env/configCheck logs --previous_
PendingInsufficient resources, no node matchesCheck events, node capacity_
OOMKilledMemory limit exceededIncrease limit or fix leak_
CreateContainerConfigErrorMissing ConfigMap/SecretCheck referenced resources exist
Evicted_Node disk/memory pressureClean up node, increase resources_
Back-off restartingReadiness probe failingCheck probe config, port, path_
connection refusedService not ready, wrong portCheck endpoints, service ports_
DNS resolution failedCoreDNS down, wrong service name_Check coredns pods, FQDN_

💡 KEY TAKEAWAYS

  1. Systematic approach: Identify → Isolate → Diagnose → Fix → Verify
  2. kubectl describe: First tool for any K8s issue
  3. Events: kubectl get events --sort-by=.lastTimestamp
  4. Debug containers: Ephemeral containers for live diagnosis
  5. Layer-by-layer: App → Container → Pod → Service → Node → Cluster
  6. Document: Every fix becomes a runbook entry

🎯 EXERCISE

Exercise 1: Troubleshooting Challenge

  • Teammate creates 5 broken deployments (wrong image, missing configmap, etc.)
  • Diagnose and fix each using only kubectl__HTMLTAG_225___
  • Document diagnostic steps for each

Exercise 2: Diagnostic Toolkit

  • Create a "debug toolbox" pod with netshoot, pg_isready, redis-cli
  • Practice DNS, network, storage troubleshooting
  • Build personal troubleshooting cheat sheet

📚 NEXT POST

In Lesson 50: Capstone Project — E-Commerce Platform, we will build and deploy the entire end-to-end microservices system.