Chuyển đến nội dung chính

Lesson 15: Production Deployment — Case Study sending 10 million emails

End-to-end case study: design and implementation of a system to send 10 million emails for marketing campaign. Infrastructure setup, Kubernetes deployment, CI/CD, load testing, chaos scenarios, cost analysis and production lessons learned.

🏗️ Architecture — Lesson 15 Lesson 15: Production Deployment — Case Study sent 10 million emails

Design a Notification System to send millions of Emails

Part 5: Deliverability, Monitoring & Production

xdev.asia

Introduction

We end the series with a realistic problem: an e-commerce company needs to send 10 million flash sale emails in 4 hours, while still keeping the transactional flow running normally. This article combines all the previous pieces into a production-ready design.


1. Problem and input assumptions

Business requirements

  • Send 10 million emails in up to 4 hours.
  • There is basic personalization by name, discount code, locale.
  • There are unsubscribe, open tracking, click tracking.
  • Does not affect OTP and order confirmation.

Capacity target

10,000,000 / 14,400 giây ≈ 694 emails/giây
Peak headroom x2 -> thiết kế cho ~1,400 emails/giây

Operational architecture

  • 1 main provider: Amazon SES.
  • 1 backup provider: SendGrid.
  • 2 separate worker pools: transactional and bulk marketing.
  • 1 Redis cluster for rate limiting and scheduling.
  • 1 Kafka cluster as event-driven backbone.
  • 1 PostgreSQL primary + read replicas for metadata and analytics ingestion.

2. Overall production architecture

Admin UI / Campaign API
        │
        ▼
Campaign Planner
        │
        ├── Recipient Snapshot Service
        ├── Batch Planner
        └── Kafka topics
                │
                ▼
          Bulk Worker Pool
                │
        ┌───────┴────────┐
        ▼                ▼
   Amazon SES        SendGrid Fallback
        │                │
        └───────┬────────┘
                ▼
        Webhook Ingestion
                │
                ▼
         Status Aggregator
                │
                ▼
      PostgreSQL + Grafana/Prometheus

Main ingredients

IngredientsRole
Campaign Plannercreate snapshots and batch jobs
Kafkadecouple producer/consumer
Redislimiter, retry schedule, distributed locks
Bulk Workersrender + send campaign traffic
Transactional WorkersEnsure critical traffic
Webhook Ingestionreceive delivery/bounce/complaint events

3. Recommended Kubernetes infrastructure

Separate workloads by namespaces and deployments

apiVersion: apps/v1
kind: Deployment
metadata:
  name: bulk-email-workers
spec:
  replicas: 12
  selector:
    matchLabels:
      app: bulk-email-workers
  template:
    metadata:
      labels:
        app: bulk-email-workers
    spec:
      containers:
        - name: worker
          image: ghcr.io/xdev/notification-workers:2026.04.01
          env:
            - name: WORKER_GROUP
              value: bulk
            - name: KAFKA_CONSUMER_GROUP
              value: bulk-workers
          resources:
            requests:
              cpu: "500m"
              memory: "512Mi"
            limits:
              cpu: "2"
              memory: "2Gi"
          readinessProbe:
            httpGet:
              path: /health/ready
              port: 8080

Initial sizing suggestion

ServiceQuantityNotes
API / Planner3 podsBasic BP
Bulk Workers12-40 podsautoscale by lag
Transactional Workers4-8 podsreserved capacity
Webhook Processors3-6 podsscale by callback burst
Redis3 nodessentinel/cluster
Kafka3 brokersreplication factor 3

4. CI/CD and release strategy

Pipeline should be there

  1. Unit tests for template rendering, limiter, provider adapters.
  2. Integration tests with Kafka, Redis, PostgreSQL.
  3. Smoke test sends email via sandbox provider.
  4. Canary deploy to the new worker version.
  5. Rollback quickly if send failure rate increases.

Why do workers need canary?

One small bug in the renderer or provider adapter can turn 10 million emails into 10 million errors. Canary 1-5% traffic helps detect regressions before the campaign is widely affected.


5. Load testing with k6

It is impossible to test production traffic without load testing the campaign coordination part.

k6 example for Campaign API

import http from 'k6/http';
import { check, sleep } from 'k6';

export const options = {
  scenarios: {
    create_campaigns: {
      executor: 'ramping-vus',
      startVUs: 1,
      stages: [
        { duration: '1m', target: 20 },
        { duration: '3m', target: 100 },
        { duration: '2m', target: 0 },
      ],
    },
  },
};

export default function () {
  const payload = JSON.stringify({
    campaign_id: `camp-${__VU}-${__ITER}`,
    template_id: 'flash_sale_v2',
    segment_id: 'active_users_30d',
    priority: 'normal',
  });

  const response = http.post('https://api.example.com/campaigns', payload, {
    headers: { 'Content-Type': 'application/json' },
  });

  check(response, {
    'campaign accepted': (r) => r.status === 202,
  });

  sleep(1);
}

What to test besides the API

  • Queue backlog growth when provider throttle.
  • Worker autoscaling when lag increases.
  • Redis limits latency when concurrent access is high.
  • Webhook ingestion burst after provider flush events.

6. Chaos scenarios should be simulated

SituationExpectations
SES pay 429 lasts 15 minutesdeceleration + partial failover to SendGrid
Redis increases latencyworker degrades but does not duplicate massively
20% of worker pods killedbatch leases are reclaimed and resume
Webhook processor downtimeevents are buffered, no state is lost
Template bug on a campaigncampaign is paused, other traffic is still safe

The goal of chaos testing

Not to prove the system is immortal, but to confirm that when it fails it fails in a controlled, observable, and recoverable way.


7. Preliminary cost analysis

Large cost component

CategoryEstimate
Amazon SES sends 10M emailsabout $1,000
SendGrid fallback reservea few hundred to a few thousand USD depending on the plan
Kubernetes computedepends on cloud and autoscale window
Kafka/Redis/PostgreSQLfixed base costs
ObservabilityPrometheus/Grafana managed or self-hosted

Cost optimization

  • Use SES as the main provider for high-volume workloads.
  • Only enable the fallback provider at a level sufficient for disaster scenarios.
  • Separate heavy analytics into an async pipeline, without forcing main PostgreSQL to shoulder the entire load.
  • Optimize template rendering cache to reduce worker CPU.

8. Lessons learned from production

The right decisions

  • Separate transactional and bulk workers from the beginning.
  • Identifies the message with message_id Stable to idempotent.
  • Build a dashboard campaign ETA and complaint rate before running a large campaign.
  • Warm-up domain/IP more carefully than initially expected.

Painful but valuable lessons

  1. The theoretical throughput of the worker is not as important as the actual throughput through the provider.
  2. A bad marketing campaign can damage the reputation of even transaction traffic if using the same domain/IP.
  3. Retry without jitter will quickly turn into self-inflicted DDoS.
  4. Webhook reconciliation is required to know which emails are actually delivered.

9. Go-live checklist for 10 million email campaign

  • SPF, DKIM, DMARC domains have been authenticated and aligned.
  • Segment has been cleaned, suppression list applied.
  • Rate limits according to configured provider/domain/IP.
  • Dashboard, alerts and runbooks are ready.
  • Fallback provider has been tested.
  • The small Canary campaign ran successfully.
  • On-call rotation knows exactly the page threshold and how to pause the campaign.

Summary

The problem of sending 10 million emails is not just a scale worker. It is a simultaneous problem of event-driven architecture, deliverability, rate control, system monitoring and operating procedures. When these layers are designed together, a large campaign becomes a predictable and controllable workload instead of a gamble.

You have gone through the core knowledge chain to design a large-scale email notification platform, from high-level design to production deployment.