Chuyển đến nội dung chính

BÀI 46: PRODUCTION READINESS CHECKLIST

Production readiness review toàn diện: infrastructure, security, observability, reliability, performance, compliance checklist, và go-live planning.

🔒 DevSecOps — Bài 46 BÀI 46: PRODUCTION READINESS CHECKLIST

Deploy Microservices On-Premises với Kubernetes HA

Phần 12: Production Operations & Capstone Project

xdev.asia

🎯 MỤC TIÊU BÀI HỌC

  • ✅ Production readiness review framework
  • ✅ Infrastructure checklist
  • ✅ Security hardening checklist
  • ✅ Observability & reliability checklist
  • ✅ Go-live planning và change management

PHẦN 1: INFRASTRUCTURE CHECKLIST

#CategoryItemStatus
1K8s Cluster3+ control plane nodes (HA)☐
2K8s Cluster3+ worker nodes (anti-affinity)☐
3K8s Clusteretcd backup scheduled (hourly)☐
4K8s ClusterKubernetes version current (N-1)☐
5NetworkingCNI installed (Cilium) + NetworkPolicies☐
6NetworkingMetalLB LoadBalancer configured☐
7NetworkingIstio service mesh + mTLS☐
8StorageRook-Ceph cluster healthy (3+ OSDs)☐
9StorageStorageClass default set☐
10StorageVolumeSnapshot class configured☐
11DatabasePostgreSQL HA (3 replicas, sync replication)☐
12DatabaseAutomated backup + PITR tested☐
13DatabaseConnection pooling (PgBouncer)☐
14MQRabbitMQ/Kafka cluster HA☐
15CacheRedis Sentinel/Cluster HA☐

PHẦN 2: SECURITY CHECKLIST

#ItemStatus
1RBAC: No cluster-admin for applications☐
2Pod Security Standards: Restricted enforced☐
3ServiceAccount: Auto-mount disabled☐
4Secrets: Stored in Vault (not plain K8s secrets)☐
5Network Policies: Default deny-all per namespace☐
6Kyverno: Validation policies enforced☐
7Falco: Runtime security monitoring active☐
8Harbor: Images scanned, no critical CVEs☐
9Image signing: cosign verification enabled☐
10Audit logging: Enabled, forwarded to Loki☐
11etcd encryption at rest enabled☐
12TLS everywhere (Istio mTLS + ingress TLS)☐

PHẦN 3: OBSERVABILITY CHECKLIST

#ItemStatus
1Prometheus: Metrics collection for all services☐
2Loki: Centralized logging with structured JSON☐
3Tempo: Distributed tracing with OTel☐
4Grafana: 3-level dashboards (platform → service → request)☐
5Correlation: Trace-Log-Metric linking configured☐
6SLOs defined: Availability + Latency per service☐
7Alerting: Multi-burn-rate SLO alerts☐
8Alert routing: Critical → PagerDuty, Warning → Slack☐
9On-call rotation configured☐
10Runbooks linked to alerts☐

PHẦN 4: RELIABILITY CHECKLIST

#ItemStatus
1HPA configured for stateless services☐
2PodDisruptionBudget for all critical workloads☐
3Liveness + readiness probes on all containers☐
4Resource requests + limits set on all pods☐
5Pod anti-affinity: spread across nodes☐
6Circuit breaker configured (Istio DestinationRule)☐
7Retry + timeout policies in VirtualService☐
8Velero backup tested (restore verified)☐
9DR runbook documented + tested☐
10Chaos engineering: GameDay completed☐

PHẦN 5: GO-LIVE PLANNING


Go-Live Timeline:

T-2 weeks: Feature freeze, final testing
T-1 week:  Performance testing, DR drill, security scan
T-3 days:  Staging deployment with production data clone
T-1 day:   Final review meeting, rollback plan confirmed
T-0:       Go-live (off-peak hours)

Go-Live Day:
 08:00  Pre-checks (all systems green)
 09:00  DNS cutover / traffic shift
 09:30  Smoke tests
 10:00  Gradual traffic ramp (10% → 25% → 50% → 100%)
 12:00  Full traffic
 18:00  Post-launch review
 
Rollback Plan:
 - DNS revert to old infrastructure
 - Estimated rollback time: 5 minutes

💡 KEY TAKEAWAYS

  1. Checklist: Systematic review prevents "forgot to configure X"
  2. Categories: Infrastructure, Security, Observability, Reliability
  3. Go-live: Gradual traffic ramp, always have rollback plan
  4. Review: Peer review checklist before production
  5. Living document: Update checklist after each incident

🎯 BÀI TẬP

Bài tập 1: Readiness Review

  • Run through all checklists for your cluster
  • Document gaps and create remediation plan
  • Perform peer review with teammate

📚 BÀI TIẾP THEO

Trong Bài 47: Day-2 Operations & Maintenance, chúng ta sẽ học vận hành production hàng ngày.