Chuyển đến nội dung chính

Lesson 23: Deployment Strategies — Canary, Blue/Green & Progressive Delivery

Rolling Update, Blue/Green Deployment, Canary Release, A/B Testing, progressive delivery with Argo Rollouts/Flagger, automated canary analysis, rollback strategies and feature flags.

🏗️ Architecture — Lesson 23 Lesson 23: Deployment Strategies — Canary, Blue/Green & Progressive Delivery

Cloud Native Microservices Architecture

Part 7: CI/CD & Deployment Strategies

xdev.asia

Lesson 23: Deployment Strategies — Canary, Blue/Green & Progressive Delivery

Introduction

Deploying a new version to production always poses risks. The question is: How to deploy to minimize risk?

Modern deployment strategies let you roll out changes gradually, observe behavior, and roll back immediately if there are problems — instead of "deploying everything and praying."


1. Rolling Update — Kubernetes Default

1.1 Mechanism

Before:  [v1][v1][v1][v1][v1]  (5 pods v1)

Step 1:  [v1][v1][v1][v1][v2]  (+1 v2, 0 downtime)
Step 2:  [v1][v1][v1][v2][v2]  (-1 v1, +1 v2)
Step 3:  [v1][v1][v2][v2][v2]
Step 4:  [v1][v2][v2][v2][v2]
After:   [v2][v2][v2][v2][v2]  (5 pods v2)

Traffic: Luôn có pods phục vụ (v1 hoặc v2)

1.2 Configuration

spec:
  replicas: 5
  strategy:
    type: RollingUpdate
    rollingUpdate:
      maxSurge: 1         # Tối đa 1 pod extra trong quá trình update
      maxUnavailable: 0   # Không có pod nào bị down (0 downtime)

1.3 Limitations

  • During deployment, there are both v1 and v2 serving traffic
  • Not suitable when v2 has breaking API changes compared to v1
  • Difficult to rollback quickly (have to wait for rolling update to reverse)

2. Blue/Green Deployment

2.1 Mechanism

Blue environment (v1 — currently live):
  [v1][v1][v1][v1][v1]
        ↑
    Traffic (100%)

Chuẩn bị Green environment (v2 — idle):
  [v2][v2][v2][v2][v2]
        ↑
    Traffic (0%) — chỉ internal testing

Switch traffic (atomic):
  Blue  [v1][v1][v1][v1][v1] ← Traffic (0%)
  Green [v2][v2][v2][v2][v2] ← Traffic (100%)

Nếu v2 OK: Hủy Blue environment
Nếu v2 lỗi: Switch lại Blue (< 1 phút)

2.2 Kubernetes Implementation

# Blue Service (v1)
apiVersion: v1
kind: Service
metadata:
  name: order-service
spec:
  selector:
    app: order-service
    version: blue         # ← Chỉ route vào Blue pods
  ports:
    - port: 8080

---
# Blue Deployment
apiVersion: apps/v1
kind: Deployment
metadata:
  name: order-service-blue
spec:
  replicas: 5
  selector:
    matchLabels:
      app: order-service
      version: blue
  template:
    metadata:
      labels:
        app: order-service
        version: blue
    spec:
      containers:
        - name: app
          image: order-service:v1

---
# Green Deployment (running idle, being tested)
apiVersion: apps/v1
kind: Deployment
metadata:
  name: order-service-green
spec:
  replicas: 5
  selector:
    matchLabels:
      app: order-service
      version: green
  template:
    metadata:
      labels:
        app: order-service
        version: green
    spec:
      containers:
        - name: app
          image: order-service:v2
# Switch traffic sang Green (atomic patch)
kubectl patch service order-service -p '{"spec":{"selector":{"version":"green"}}}'

# Rollback — switch lại Blue (< 5 giây)
kubectl patch service order-service -p '{"spec":{"selector":{"version":"blue"}}}'

2.3 Advantages and disadvantages

Advantages: Zero-downtime, extremely fast rollback, test v2 before going live

Disadvantages: Consumes double resources (running both Blue and Green), session state must be stateless (because the switch will lose the session)


3. Canary Release

3.1 Mechanism

Instead of switching everything, send a small amount of traffic to the new version first:

  v1 pods (90% traffic):  [v1][v1][v1][v1][v1]
  v2 pods (10% traffic):  [v2]

  Sau 30 phút, không có vấn đề → tăng:
  v1 pods (75% traffic):  [v1][v1][v1]
  v2 pods (25% traffic):  [v2][v2]

  Sau 1 giờ → tăng:
  v1 pods (50%) / v2 pods (50%)

  Sau 2 giờ → v2 = 100%

3.2 Istio Traffic Splitting

# VirtualService — định nghĩa traffic split
apiVersion: networking.istio.io/v1beta1
kind: VirtualService
metadata:
  name: order-service
spec:
  hosts:
    - order-service
  http:
    - route:
        - destination:
            host: order-service
            subset: v1
          weight: 90       # 90% traffic → v1
        - destination:
            host: order-service
            subset: v2
          weight: 10       # 10% traffic → v2

---
# DestinationRule — định nghĩa subsets
apiVersion: networking.istio.io/v1beta1
kind: DestinationRule
metadata:
  name: order-service
spec:
  host: order-service
  subsets:
    - name: v1
      labels:
        version: v1
    - name: v2
      labels:
        version: v2

4. Progressive Delivery with Argo Rollouts

Argo Rollouts automates canary deployment with automated analysis:

4.1 Rollout Definition

apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
  name: order-service
spec:
  replicas: 10
  selector:
    matchLabels:
      app: order-service

  template:
    metadata:
      labels:
        app: order-service
    spec:
      containers:
        - name: order-service
          image: order-service:v2  # ← CI update này

  strategy:
    canary:
      # Các bước canary tự động
      steps:
        - setWeight: 10   # Bước 1: 10% traffic
        - pause: {duration: 10m}   # Chờ 10 phút
        - analysis:
            templates:
              - templateName: success-rate
        - setWeight: 25   # Bước 2: 25% nếu analysis pass
        - pause: {duration: 10m}
        - setWeight: 50   # Bước 3: 50%
        - pause: {duration: 10m}
        - setWeight: 100  # Bước 4: 100%

      # Service Mesh integration
      trafficRouting:
        istio:
          virtualService:
            name: order-service

      # Nếu analysis fail → tự động rollback
      autoPromotionEnabled: false  # Require explicit promotion

4.2 AnalysisTemplate — Automated Canary Analysis

apiVersion: argoproj.io/v1alpha1
kind: AnalysisTemplate
metadata:
  name: success-rate
spec:
  args:
    - name: service-name
  metrics:
    # Metric 1: Error rate phải < 5%
    - name: success-rate
      interval: 1m
      count: 5          # Đo 5 lần
      successCondition: result[0] >= 0.95
      failureLimit: 2   # Cho phép fail ≤ 2 lần
      provider:
        prometheus:
          address: http://prometheus.monitoring:9090
          query: |
            sum(rate(http_requests_total{
              service="{{args.service-name}}",
              status!~"5.."
            }[5m]))
            /
            sum(rate(http_requests_total{
              service="{{args.service-name}}"
            }[5m]))

    # Metric 2: p99 latency phải < 500ms
    - name: latency-p99
      interval: 1m
      count: 5
      successCondition: result[0] < 0.5
      failureLimit: 2
      provider:
        prometheus:
          address: http://prometheus.monitoring:9090
          query: |
            histogram_quantile(0.99,
              sum(rate(http_request_duration_seconds_bucket{
                service="{{args.service-name}}"
              }[5m])) by (le)
            )

    # Metric 3: Compare canary vs baseline (kayenta style)
    - name: error-rate-canary-vs-stable
      interval: 2m
      count: 3
      successCondition: result[0] <= 1.5  # Canary không tệ hơn stable 1.5x
      provider:
        prometheus:
          query: |
            (
              sum(rate(http_requests_total{service="order-service",version="canary",status=~"5.."}[5m]))
              /
              sum(rate(http_requests_total{service="order-service",version="canary"}[5m]))
            ) / (
              sum(rate(http_requests_total{service="order-service",version="stable",status=~"5.."}[5m]))
              /
              sum(rate(http_requests_total{service="order-service",version="stable"}[5m]))
            )

4.3 Manual Control

# CLI để interact với rollout
kubectl argo rollouts get rollout order-service --watch

# Promote manually (nếu autoPromotionEnabled: false)
kubectl argo rollouts promote order-service

# Abort và rollback
kubectl argo rollouts abort order-service

# Undo về stable
kubectl argo rollouts undo order-service

5. A/B Testing

A/B Testing route traffic based on user attributes (not random %):

# Istio VirtualService với header-based routing
apiVersion: networking.istio.io/v1beta1
kind: VirtualService
metadata:
  name: order-service
spec:
  http:
    # Route users với header X-Beta-User: true → v2
    - match:
        - headers:
            x-beta-user:
              exact: "true"
      route:
        - destination:
            host: order-service
            subset: v2

    # Route users ở EU → v2 (geographic test)
    - match:
        - headers:
            x-user-region:
              exact: "EU"
      route:
        - destination:
            host: order-service
            subset: v2

    # Tất cả còn lại → v1
    - route:
        - destination:
            host: order-service
            subset: v1

6. Feature Flags

Feature flags separate deploy and release:

Deploy: Upload code lên server (mọi user không thấy)
Release: Bật feature cho users chọn lọc (flag = on)

Cho phép:
- Dark launch: deploy trước, release sau
- Ring deployment: internal → beta users → 10% → 100%
- Hotfix without redeploy: tắt feature ngay bằng flag
// OpenFeature + FlagD
@Autowired
Client featureClient;

public Order createOrder(CreateOrderRequest request) {
    // Check feature flag
    boolean useNewCheckout = featureClient.getBooleanValue(
        "new-checkout-flow",
        false,  // default
        EvaluationContext.builder()
            .targetingKey(request.getCustomerId())
            .add("plan", request.getCustomer().getPlan())
            .build()
    );

    if (useNewCheckout) {
        return newCheckoutService.process(request);
    } else {
        return legacyCheckoutService.process(request);
    }
}
# FlagD manifest (feature flag definitions)
flags:
  new-checkout-flow:
    state: ENABLED
    variants:
      "on": true
      "off": false
    defaultVariant: "off"
    targeting:
      # 20% users ngẫu nhiên
      fractional:
        - ["on", 20]
        - ["off", 80]

7. Compare strategies

StrategyDowntimeResource CostRollback Speed ​​Risk LevelSuitable
Rolling Update01xSlow (rolling)AverageDefault for stateless services
Blue/Green02xFast (<1 min)LowWhen you need to rollback immediately
Canaries01xFastVery lowWhen you want to validate on real traffic
A/B Testing01xFastLowValidate UX/business metrics
Feature Flags01xInstant (flag off)Very lowGradual release, dark launch

8. Best Practices

1. Luôn có readiness probe trước khi traffic vào
2. Graceful shutdown — drain in-flight requests trước khi pod down
3. Không deploy vào peak hours
4. Automated rollback khi error rate tăng
5. Canary trước production cho mọi breaking change
6. Feature flags cho long-running features
7. Post-deployment monitoring 30 phút
8. Runbook sẵn sàng trước mỗi deploy

Summary

ConceptPurpose
Rolling UpdateZero-downtime, default Kubernetes strategy
Blue/GreenInstant rollback, no traffic mixing
CanariesValidate on small % real traffic
Argo RolloutsAutomated progressive delivery with analysis
AnalysisTemplateAutomatically pass/fail canary based on metrics
A/B TestingRoute according to user attributes
Feature FlagsSeparate deploy and release

Next article: Authentication & Authorization — OAuth2, JWT & OIDC