Chuyển đến nội dung chính

BÀI 47: DAY-2 OPERATIONS & MAINTENANCE

Kubernetes cluster upgrades, node maintenance, certificate rotation, capacity planning, incident management, on-call practices, và operational runbooks.

🔒 DevSecOps — Bài 47 BÀI 47: DAY-2 OPERATIONS & MAINTENANCE

Deploy Microservices On-Premises với Kubernetes HA

Phần 12: Production Operations & Capstone Project

xdev.asia

🎯 MỤC TIÊU BÀI HỌC

  • ✅ Kubernetes version upgrade process
  • ✅ Node maintenance (drain, cordon, uncordon)
  • ✅ Certificate rotation
  • ✅ Capacity planning và trending
  • ✅ Incident management framework
  • ✅ Operational runbooks

PHẦN 1: KUBERNETES UPGRADE


Upgrade Strategy (v1.30 → v1.31):

Order: Control Plane first, then Workers

Step 1: Upgrade control plane nodes (one at a time)
  master-1 → master-2 → master-3

Step 2: Upgrade worker nodes (rolling)
  worker-1 → worker-2 → worker-3 → worker-4

Each node:
  cordon → drain → upgrade → uncordon
# Upgrade control plane (master-1):
# 1. Update kubeadm:
apt-get update
apt-get install -y kubeadm=1.31.0-1.1

# 2. Plan upgrade:
kubeadm upgrade plan

# 3. Apply upgrade:
kubeadm upgrade apply v1.31.0

# 4. Drain node:
kubectl drain master-1 --ignore-daemonsets --delete-emptydir-data

# 5. Upgrade kubelet + kubectl:
apt-get install -y kubelet=1.31.0-1.1 kubectl=1.31.0-1.1
systemctl daemon-reload
systemctl restart kubelet

# 6. Uncordon:
kubectl uncordon master-1

# Upgrade worker nodes:
# 1. Drain:
kubectl drain worker-1 --ignore-daemonsets --delete-emptydir-data

# 2. SSH to worker, upgrade:
apt-get update
apt-get install -y kubeadm=1.31.0-1.1
kubeadm upgrade node
apt-get install -y kubelet=1.31.0-1.1
systemctl daemon-reload && systemctl restart kubelet

# 3. Uncordon:
kubectl uncordon worker-1

PHẦN 2: NODE MAINTENANCE

# Planned maintenance (e.g., hardware replacement):

# 1. Cordon (prevent new pods):
kubectl cordon worker-02

# 2. Drain (evict existing pods):
kubectl drain worker-02 \
  --ignore-daemonsets \
  --delete-emptydir-data \
  --grace-period=120 \
  --timeout=300s

# 3. Perform maintenance (reboot, disk replace, etc.)

# 4. Uncordon (allow pods again):
kubectl uncordon worker-02

# 5. Verify:
kubectl get nodes
kubectl get pods -o wide | grep worker-02
# PodDisruptionBudget (protect during drain):
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
  name: order-service-pdb
spec:
  minAvailable: 2
  # or: maxUnavailable: 1
  selector:
    matchLabels:
      app: order-service

PHẦN 3: CERTIFICATE ROTATION

# Check certificate expiry:
kubeadm certs check-expiration

# Output:
# CERTIFICATE                EXPIRES           RESIDUAL TIME
# admin.conf                 Jun 15, 2025      364d
# apiserver                  Jun 15, 2025      364d
# apiserver-etcd-client      Jun 15, 2025      364d
# apiserver-kubelet-client   Jun 15, 2025      364d
# controller-manager.conf    Jun 15, 2025      364d
# etcd-healthcheck-client    Jun 15, 2025      364d
# etcd-peer                  Jun 15, 2025      364d
# etcd-server                Jun 15, 2025      364d
# front-proxy-client         Jun 15, 2025      364d
# scheduler.conf             Jun 15, 2025      364d

# Renew all certificates:
kubeadm certs renew all

# Restart control plane pods:
kubectl -n kube-system delete pod -l tier=control-plane

# Update kubeconfig:
cp /etc/kubernetes/admin.conf ~/.kube/config
# Alert on certificate expiry:
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
  name: cert-expiry-alerts
spec:
  groups:
    - name: certificates
      rules:
        - alert: KubeCertExpiringSoon
          expr: |
            apiserver_client_certificate_expiration_seconds_count > 0
            and
            apiserver_client_certificate_expiration_seconds_bucket{le="604800"} > 0
          labels:
            severity: warning
          annotations:
            summary: "K8s certificate expiring within 7 days"

PHẦN 4: CAPACITY PLANNING

# Grafana queries for capacity trending:

# CPU usage trend (predict when 80% reached):
predict_linear(
  sum(rate(node_cpu_seconds_total{mode!="idle"}[1h])) by (instance)
  [7d:1h], 30*86400
)

# Memory usage trend:
predict_linear(
  node_memory_MemAvailable_bytes[7d:1h], 30*86400
)

# Disk usage trend:
predict_linear(
  node_filesystem_avail_bytes{mountpoint="/"}[7d:1h], 30*86400
)

# Pod count trending:
predict_linear(
  sum(kube_pod_info)[7d:1h], 30*86400
)
MetricThresholdAction
Cluster CPU allocation> 70%Plan new worker nodes
Cluster memory allocation> 75%Plan new worker nodes
Ceph storage used> 70%Add OSDs or disks
PV usage> 80%Expand PVC or add storage
Pod count vs quota> 80%Increase ResourceQuota

PHẦN 5: INCIDENT MANAGEMENT


Incident Severity Levels:

SEV1 (Critical):
  - Service completely down
  - Data loss risk
  - Response: Immediate, all-hands
  - Communication: Every 15 min
  
SEV2 (Major):
  - Service degraded (high error rate, slow)
  - Response: < 15 min
  - Communication: Every 30 min

SEV3 (Minor):
  - Non-critical component issue
  - Response: < 1 hour
  - Communication: Status update

SEV4 (Low):
  - Cosmetic, minor bug
  - Response: Next business day

Incident Flow:
Alert → Acknowledge → Triage → Mitigate → Root Cause → Postmortem
# Incident response runbook template:

# 1. SERVICE: [Service Name]
# 2. ALERT: [Alert Name]
# 3. SYMPTOMS: [What user/system sees]
# 4. DIAGNOSIS:
#    - Check pod status: kubectl get pods -l app=X
#    - Check logs: kubectl logs deploy/X --tail=50
#    - Check metrics: Grafana dashboard [URL]
#    - Check recent changes: kubectl rollout history deploy/X
# 5. MITIGATION:
#    - Rollback: kubectl rollout undo deploy/X
#    - Scale up: kubectl scale deploy/X --replicas=10
#    - Restart: kubectl rollout restart deploy/X
# 6. ESCALATION:
#    - Team lead: [name]
#    - SRE: [name]
#    - Manager: [name]

💡 KEY TAKEAWAYS

  1. Upgrades: Control plane first, workers rolling, one at a time
  2. PDB: Always set PodDisruptionBudget before drain
  3. Certificates: Monitor expiry, renew before 30-day warning
  4. Capacity: predict_linear for proactive planning
  5. Incidents: SEV levels, structured runbooks, postmortems

🎯 BÀI TẬP

Bài tập 1: Upgrade Drill

  • Perform K8s minor version upgrade on lab cluster
  • Practice node drain with PDB protection
  • Renew certificates, verify cluster health

Bài tập 2: Incident Response

  • Create runbook for top 5 common alerts
  • Practice incident simulation: inject failure → follow runbook
  • Write postmortem template

📚 BÀI TIẾP THEO

Trong Bài 48: Performance Testing & Optimization, chúng ta sẽ load test và optimize hệ thống.