Production readiness review toàn diện: infrastructure, security, observability, reliability, performance, compliance checklist, và go-live planning.
Deploy Microservices On-Premises với Kubernetes HA
Phần 12: Production Operations & Capstone Project
xdev.asia
🎯 MỤC TIÊU BÀI HỌC
- ✅ Production readiness review framework
- ✅ Infrastructure checklist
- ✅ Security hardening checklist
- ✅ Observability & reliability checklist
- ✅ Go-live planning và change management
PHẦN 1: INFRASTRUCTURE CHECKLIST
| # | Category | Item | Status |
| 1 | K8s Cluster | 3+ control plane nodes (HA) | ☐ |
| 2 | K8s Cluster | 3+ worker nodes (anti-affinity) | ☐ |
| 3 | K8s Cluster | etcd backup scheduled (hourly) | ☐ |
| 4 | K8s Cluster | Kubernetes version current (N-1) | ☐ |
| 5 | Networking | CNI installed (Cilium) + NetworkPolicies | ☐ |
| 6 | Networking | MetalLB LoadBalancer configured | ☐ |
| 7 | Networking | Istio service mesh + mTLS | ☐ |
| 8 | Storage | Rook-Ceph cluster healthy (3+ OSDs) | ☐ |
| 9 | Storage | StorageClass default set | ☐ |
| 10 | Storage | VolumeSnapshot class configured | ☐ |
| 11 | Database | PostgreSQL HA (3 replicas, sync replication) | ☐ |
| 12 | Database | Automated backup + PITR tested | ☐ |
| 13 | Database | Connection pooling (PgBouncer) | ☐ |
| 14 | MQ | RabbitMQ/Kafka cluster HA | ☐ |
| 15 | Cache | Redis Sentinel/Cluster HA | ☐ |
PHẦN 2: SECURITY CHECKLIST
| # | Item | Status |
| 1 | RBAC: No cluster-admin for applications | ☐ |
| 2 | Pod Security Standards: Restricted enforced | ☐ |
| 3 | ServiceAccount: Auto-mount disabled | ☐ |
| 4 | Secrets: Stored in Vault (not plain K8s secrets) | ☐ |
| 5 | Network Policies: Default deny-all per namespace | ☐ |
| 6 | Kyverno: Validation policies enforced | ☐ |
| 7 | Falco: Runtime security monitoring active | ☐ |
| 8 | Harbor: Images scanned, no critical CVEs | ☐ |
| 9 | Image signing: cosign verification enabled | ☐ |
| 10 | Audit logging: Enabled, forwarded to Loki | ☐ |
| 11 | etcd encryption at rest enabled | ☐ |
| 12 | TLS everywhere (Istio mTLS + ingress TLS) | ☐ |
PHẦN 3: OBSERVABILITY CHECKLIST
| # | Item | Status |
| 1 | Prometheus: Metrics collection for all services | ☐ |
| 2 | Loki: Centralized logging with structured JSON | ☐ |
| 3 | Tempo: Distributed tracing with OTel | ☐ |
| 4 | Grafana: 3-level dashboards (platform → service → request) | ☐ |
| 5 | Correlation: Trace-Log-Metric linking configured | ☐ |
| 6 | SLOs defined: Availability + Latency per service | ☐ |
| 7 | Alerting: Multi-burn-rate SLO alerts | ☐ |
| 8 | Alert routing: Critical → PagerDuty, Warning → Slack | ☐ |
| 9 | On-call rotation configured | ☐ |
| 10 | Runbooks linked to alerts | ☐ |
PHẦN 4: RELIABILITY CHECKLIST
| # | Item | Status |
| 1 | HPA configured for stateless services | ☐ |
| 2 | PodDisruptionBudget for all critical workloads | ☐ |
| 3 | Liveness + readiness probes on all containers | ☐ |
| 4 | Resource requests + limits set on all pods | ☐ |
| 5 | Pod anti-affinity: spread across nodes | ☐ |
| 6 | Circuit breaker configured (Istio DestinationRule) | ☐ |
| 7 | Retry + timeout policies in VirtualService | ☐ |
| 8 | Velero backup tested (restore verified) | ☐ |
| 9 | DR runbook documented + tested | ☐ |
| 10 | Chaos engineering: GameDay completed | ☐ |
PHẦN 5: GO-LIVE PLANNING
Go-Live Timeline:
T-2 weeks: Feature freeze, final testing
T-1 week: Performance testing, DR drill, security scan
T-3 days: Staging deployment with production data clone
T-1 day: Final review meeting, rollback plan confirmed
T-0: Go-live (off-peak hours)
Go-Live Day:
08:00 Pre-checks (all systems green)
09:00 DNS cutover / traffic shift
09:30 Smoke tests
10:00 Gradual traffic ramp (10% → 25% → 50% → 100%)
12:00 Full traffic
18:00 Post-launch review
Rollback Plan:
- DNS revert to old infrastructure
- Estimated rollback time: 5 minutes
💡 KEY TAKEAWAYS
- Checklist: Systematic review prevents "forgot to configure X"
- Categories: Infrastructure, Security, Observability, Reliability
- Go-live: Gradual traffic ramp, always have rollback plan
- Review: Peer review checklist before production
- Living document: Update checklist after each incident
🎯 BÀI TẬP
Bài tập 1: Readiness Review
- Run through all checklists for your cluster
- Document gaps and create remediation plan
- Perform peer review with teammate
📚 BÀI TIẾP THEO
Trong Bài 47: Day-2 Operations & Maintenance, chúng ta sẽ học vận hành production hàng ngày.