Chuyển đến nội dung chính

BÀI 40: CANARY & BLUE-GREEN DEPLOYMENT

Deployment strategies chi tiết: Canary với Istio/Argo Rollouts, Blue-Green deployment, traffic shifting, automated rollback, và progressive delivery.

🔒 DevSecOps — Bài 40 BÀI 40: CANARY & BLUE-GREEN DEPLOYMENT

Deploy Microservices On-Premises với Kubernetes HA

Phần 10: Deployment Patterns & Auto-Scaling

xdev.asia

🎯 MỤC TIÊU BÀI HỌC

  • ✅ So sánh deployment strategies (Rolling, Blue-Green, Canary)
  • ✅ Argo Rollouts cho progressive delivery
  • ✅ Canary deployment với traffic shifting
  • ✅ Blue-Green với instant cutover
  • ✅ Analysis templates và automated rollback

PHẦN 1: DEPLOYMENT STRATEGIES


graph LR
    subgraph ROLLING["🔄 Rolling Update — default K8s"]
        direction LR
        R1["v1 ████████"] -->|"gradually replace"| R2["v1 ██ v2 ██████"] --> R3["v2 ████████"]
    end

    subgraph BLUEGREEN["🔵🟢 Blue-Green"]
        direction LR
        BG1["🔵 Blue v1 ACTIVE
🟢 Green v2 standby"] -->|"instant switch"| BG2["🔵 Blue v1 standby
🟢 Green v2 ACTIVE"] end subgraph CANARY["🐤 Canary"] direction LR C1["v1 100%"] -->|"5%"| C2["v1 95%
v2 5%"] -->|"validate"| C3["v1 70%
v2 30%"] -->|"promote"| C4["v2 100%"] end style ROLLING fill:#1e3a5f,stroke:#3b82f6,color:#e2e8f0 style BLUEGREEN fill:#1e3a5f,stroke:#10b981,color:#e2e8f0 style CANARY fill:#1e3a5f,stroke:#f59e0b,color:#e2e8f0

stateDiagram-v2
    [*] --> Deploy_v2: New version
    Deploy_v2 --> Canary_5: Shift 5% traffic
    Canary_5 --> Analysis_1: Run analysis
    Analysis_1 --> Canary_30: ✅ Pass → Shift 30%
    Analysis_1 --> Rollback: ❌ Fail
    Canary_30 --> Analysis_2: Run analysis
    Analysis_2 --> Canary_70: ✅ Pass → Shift 70%
    Analysis_2 --> Rollback: ❌ Fail
    Canary_70 --> Full_v2: ✅ Promote 100%
    Full_v2 --> [*]: Done
    Rollback --> [*]: Reverted to v1
StrategyDowntimeRollbackRiskResource Cost
Rolling UpdateZeroSlowMediumLow (+25%)
Blue-GreenZeroInstantLowHigh (2x)
CanaryZeroFastVery LowLow (+10%)

PHẦN 2: ARGO ROLLOUTS

# Install Argo Rollouts:
helm repo add argo https://argoproj.github.io/argo-helm
helm install argo-rollouts argo/argo-rollouts \
  --namespace argo-rollouts \
  --create-namespace \
  --set dashboard.enabled=true
# Canary Rollout:
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
  name: order-service
  namespace: default
spec:
  replicas: 5
  revisionHistoryLimit: 3
  selector:
    matchLabels:
      app: order-service
  template:
    metadata:
      labels:
        app: order-service
    spec:
      containers:
        - name: app
          image: harbor.local/myproject/order-service:v2.0
          ports:
            - containerPort: 8080
          resources:
            requests:
              cpu: 100m
              memory: 128Mi
  strategy:
    canary:
      canaryService: order-service-canary
      stableService: order-service-stable
      trafficRouting:
        istio:
          virtualServices:
            - name: order-service-vsvc
              routes:
                - primary
      steps:
        # Step 1: 5% traffic to canary
        - setWeight: 5
        - pause: { duration: 5m }
        # Step 2: Run analysis
        - analysis:
            templates:
              - templateName: success-rate
            args:
              - name: service-name
                value: order-service-canary
        # Step 3: 25% traffic
        - setWeight: 25
        - pause: { duration: 10m }
        # Step 4: 50% traffic
        - setWeight: 50
        - pause: { duration: 10m }
        # Step 5: 100% → promote
        - setWeight: 100

PHẦN 3: ANALYSIS TEMPLATES

# Canary analysis: check error rate and latency:
apiVersion: argoproj.io/v1alpha1
kind: AnalysisTemplate
metadata:
  name: success-rate
spec:
  args:
    - name: service-name
  metrics:
    - name: success-rate
      interval: 60s
      count: 5
      successCondition: result[0] >= 0.99
      failureLimit: 2
      provider:
        prometheus:
          address: http://prometheus-kube-prometheus-prometheus.monitoring:9090
          query: |
            sum(rate(http_server_request_duration_seconds_count{
              service="{{args.service-name}}",
              http_status_code!~"5.."
            }[2m]))
            /
            sum(rate(http_server_request_duration_seconds_count{
              service="{{args.service-name}}"
            }[2m]))

    - name: latency-p99
      interval: 60s
      count: 5
      successCondition: result[0] < 0.5
      failureLimit: 2
      provider:
        prometheus:
          address: http://prometheus-kube-prometheus-prometheus.monitoring:9090
          query: |
            histogram_quantile(0.99,
              sum(rate(http_server_request_duration_seconds_bucket{
                service="{{args.service-name}}"
              }[2m])) by (le)
            )

PHẦN 4: BLUE-GREEN DEPLOYMENT

# Blue-Green Rollout:
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
  name: payment-service
spec:
  replicas: 3
  selector:
    matchLabels:
      app: payment-service
  template:
    metadata:
      labels:
        app: payment-service
    spec:
      containers:
        - name: app
          image: harbor.local/myproject/payment-service:v2.0
  strategy:
    blueGreen:
      activeService: payment-service-active
      previewService: payment-service-preview
      autoPromotionEnabled: false
      # Pre-promotion analysis:
      prePromotionAnalysis:
        templates:
          - templateName: smoke-test
      # Scale down old version after:
      scaleDownDelaySeconds: 300
      # Anti-affinity:
      antiAffinity:
        preferredDuringSchedulingIgnoredDuringExecution:
          weight: 100
# Smoke test analysis:
apiVersion: argoproj.io/v1alpha1
kind: AnalysisTemplate
metadata:
  name: smoke-test
spec:
  metrics:
    - name: smoke-test
      count: 1
      provider:
        job:
          spec:
            template:
              spec:
                containers:
                  - name: smoke
                    image: harbor.local/myproject/smoke-test:latest
                    command: ["./run-tests.sh"]
                    env:
                      - name: TARGET_URL
                        value: "http://payment-service-preview:8080"
                restartPolicy: Never
            backoffLimit: 0

PHẦN 5: ROLLOUT OPERATIONS

# Argo Rollouts CLI:
kubectl argo rollouts get rollout order-service -w
kubectl argo rollouts status order-service

# Manual promote canary:
kubectl argo rollouts promote order-service

# Abort (rollback):
kubectl argo rollouts abort order-service

# Retry after abort:
kubectl argo rollouts retry rollout order-service

# Blue-Green: promote preview to active:
kubectl argo rollouts promote payment-service

# Dashboard:
kubectl argo rollouts dashboard -p 3100

💡 KEY TAKEAWAYS

  1. Canary: Gradual traffic shifting, metrics-driven promotion
  2. Blue-Green: Instant cutover, full preview environment
  3. Argo Rollouts: Drop-in Deployment replacement
  4. AnalysisTemplate: Automated success/failure criteria
  5. Integration: Istio traffic routing + Prometheus metrics
  6. Rollback: Automatic on analysis failure

🎯 BÀI TẬP

Bài tập 1: Canary Deployment

  • Convert Deployment to Argo Rollout
  • Configure 5% → 25% → 50% → 100% canary steps
  • Create AnalysisTemplate with success rate check

Bài tập 2: Blue-Green + Smoke Test

  • Setup Blue-Green Rollout
  • Create smoke test Job as pre-promotion analysis
  • Inject error → verify auto-rollback

📚 BÀI TIẾP THEO

Trong Bài 41: Horizontal & Vertical Pod Autoscaling, chúng ta sẽ implement auto-scaling cho workloads.