Introduction
We end the series with a realistic problem: an e-commerce company needs to send 10 million flash sale emails in 4 hours, while still keeping the transactional flow running normally. This article combines all the previous pieces into a production-ready design.
1. Problem and input assumptions
Business requirements
- Send 10 million emails in up to 4 hours.
- There is basic personalization by name, discount code, locale.
- There are unsubscribe, open tracking, click tracking.
- Does not affect OTP and order confirmation.
Capacity target
10,000,000 / 14,400 giây ≈ 694 emails/giây
Peak headroom x2 -> thiết kế cho ~1,400 emails/giây
Operational architecture
- 1 main provider: Amazon SES.
- 1 backup provider: SendGrid.
- 2 separate worker pools: transactional and bulk marketing.
- 1 Redis cluster for rate limiting and scheduling.
- 1 Kafka cluster as event-driven backbone.
- 1 PostgreSQL primary + read replicas for metadata and analytics ingestion.
2. Overall production architecture
Admin UI / Campaign API
│
▼
Campaign Planner
│
├── Recipient Snapshot Service
├── Batch Planner
└── Kafka topics
│
▼
Bulk Worker Pool
│
┌───────┴────────┐
▼ ▼
Amazon SES SendGrid Fallback
│ │
└───────┬────────┘
▼
Webhook Ingestion
│
▼
Status Aggregator
│
▼
PostgreSQL + Grafana/Prometheus
Main ingredients
| Ingredients | Role |
|---|---|
| Campaign Planner | create snapshots and batch jobs |
| Kafka | decouple producer/consumer |
| Redis | limiter, retry schedule, distributed locks |
| Bulk Workers | render + send campaign traffic |
| Transactional Workers | Ensure critical traffic |
| Webhook Ingestion | receive delivery/bounce/complaint events |
3. Recommended Kubernetes infrastructure
Separate workloads by namespaces and deployments
apiVersion: apps/v1
kind: Deployment
metadata:
name: bulk-email-workers
spec:
replicas: 12
selector:
matchLabels:
app: bulk-email-workers
template:
metadata:
labels:
app: bulk-email-workers
spec:
containers:
- name: worker
image: ghcr.io/xdev/notification-workers:2026.04.01
env:
- name: WORKER_GROUP
value: bulk
- name: KAFKA_CONSUMER_GROUP
value: bulk-workers
resources:
requests:
cpu: "500m"
memory: "512Mi"
limits:
cpu: "2"
memory: "2Gi"
readinessProbe:
httpGet:
path: /health/ready
port: 8080
Initial sizing suggestion
| Service | Quantity | Notes |
|---|---|---|
| API / Planner | 3 pods | Basic BP |
| Bulk Workers | 12-40 pods | autoscale by lag |
| Transactional Workers | 4-8 pods | reserved capacity |
| Webhook Processors | 3-6 pods | scale by callback burst |
| Redis | 3 nodes | sentinel/cluster |
| Kafka | 3 brokers | replication factor 3 |
4. CI/CD and release strategy
Pipeline should be there
- Unit tests for template rendering, limiter, provider adapters.
- Integration tests with Kafka, Redis, PostgreSQL.
- Smoke test sends email via sandbox provider.
- Canary deploy to the new worker version.
- Rollback quickly if send failure rate increases.
Why do workers need canary?
One small bug in the renderer or provider adapter can turn 10 million emails into 10 million errors. Canary 1-5% traffic helps detect regressions before the campaign is widely affected.
5. Load testing with k6
It is impossible to test production traffic without load testing the campaign coordination part.
k6 example for Campaign API
import http from 'k6/http';
import { check, sleep } from 'k6';
export const options = {
scenarios: {
create_campaigns: {
executor: 'ramping-vus',
startVUs: 1,
stages: [
{ duration: '1m', target: 20 },
{ duration: '3m', target: 100 },
{ duration: '2m', target: 0 },
],
},
},
};
export default function () {
const payload = JSON.stringify({
campaign_id: `camp-${__VU}-${__ITER}`,
template_id: 'flash_sale_v2',
segment_id: 'active_users_30d',
priority: 'normal',
});
const response = http.post('https://api.example.com/campaigns', payload, {
headers: { 'Content-Type': 'application/json' },
});
check(response, {
'campaign accepted': (r) => r.status === 202,
});
sleep(1);
}
What to test besides the API
- Queue backlog growth when provider throttle.
- Worker autoscaling when lag increases.
- Redis limits latency when concurrent access is high.
- Webhook ingestion burst after provider flush events.
6. Chaos scenarios should be simulated
| Situation | Expectations |
|---|---|
| SES pay 429 lasts 15 minutes | deceleration + partial failover to SendGrid |
| Redis increases latency | worker degrades but does not duplicate massively |
| 20% of worker pods killed | batch leases are reclaimed and resume |
| Webhook processor downtime | events are buffered, no state is lost |
| Template bug on a campaign | campaign is paused, other traffic is still safe |
The goal of chaos testing
Not to prove the system is immortal, but to confirm that when it fails it fails in a controlled, observable, and recoverable way.
7. Preliminary cost analysis
Large cost component
| Category | Estimate |
|---|---|
| Amazon SES sends 10M emails | about $1,000 |
| SendGrid fallback reserve | a few hundred to a few thousand USD depending on the plan |
| Kubernetes compute | depends on cloud and autoscale window |
| Kafka/Redis/PostgreSQL | fixed base costs |
| Observability | Prometheus/Grafana managed or self-hosted |
Cost optimization
- Use SES as the main provider for high-volume workloads.
- Only enable the fallback provider at a level sufficient for disaster scenarios.
- Separate heavy analytics into an async pipeline, without forcing main PostgreSQL to shoulder the entire load.
- Optimize template rendering cache to reduce worker CPU.
8. Lessons learned from production
The right decisions
- Separate transactional and bulk workers from the beginning.
- Identifies the message with
message_idStable to idempotent. - Build a dashboard campaign ETA and complaint rate before running a large campaign.
- Warm-up domain/IP more carefully than initially expected.
Painful but valuable lessons
- The theoretical throughput of the worker is not as important as the actual throughput through the provider.
- A bad marketing campaign can damage the reputation of even transaction traffic if using the same domain/IP.
- Retry without jitter will quickly turn into self-inflicted DDoS.
- Webhook reconciliation is required to know which emails are actually delivered.
9. Go-live checklist for 10 million email campaign
- SPF, DKIM, DMARC domains have been authenticated and aligned.
- Segment has been cleaned, suppression list applied.
- Rate limits according to configured provider/domain/IP.
- Dashboard, alerts and runbooks are ready.
- Fallback provider has been tested.
- The small Canary campaign ran successfully.
- On-call rotation knows exactly the page threshold and how to pause the campaign.
Summary
The problem of sending 10 million emails is not just a scale worker. It is a simultaneous problem of event-driven architecture, deliverability, rate control, system monitoring and operating procedures. When these layers are designed together, a large campaign becomes a predictable and controllable workload instead of a gamble.
You have gone through the core knowledge chain to design a large-scale email notification platform, from high-level design to production deployment.