Introduction
Without metrics, a 10 million email campaign is just a leap of faith into the dark. The notification production system must be able to immediately answer the questions: how fast is it sending, where is it stuck, which provider is failing, which campaign is about to break its SLA.
1. Observe the system layer by layer
Four main layers of metrics
| Class | For example metric |
|---|---|
| Business | emails sent, delivered, complaints, campaign ETA |
| Application | render latency, send latency, retry count |
| Queue | consumer lag, queue depth, DLQ size |
| Infrastructure | CPU, memory, network, pod restarts |
Anti-patterns should be avoided
- Only look at CPU/RAM, not queue lag.
- Just look
sent_countwithout looking at delivery/bounce/complaint. - Do not attach metrics
provider,campaign_type,recipient_domain.
2. Core set of metrics for email platform
Business metrics
| Metrics | Meaning |
|---|---|
emails_requested_total | total number of emails requested to be sent |
emails_sent_total | number of emails accepted by provider |
emails_delivered_total | delivery confirmed email number |
emails_bounced_total | bounce by type |
emails_complained_total | complaint spam |
campaign_eta_seconds | Campaign completion ETA |
Application metrics
| Metrics | Labels should have |
|---|---|
template_render_seconds | template_id, locale |
provider_send_seconds | provider, priority |
retry_attempt_total | error_class, provider |
throttle_decision_total | limit_type, result |
worker_batch_duration_seconds | worker_group |
Queue metrics
| Metrics | Meaning |
|---|---|
queue_depth | total pending jobs |
consumer_lag | backlog level compared to producer |
retry_queue_depth | number of jobs waiting for retry |
dlq_messages_total | quantity goes into DLQ |
3. Prometheus instrumentation
Example metric definitions
from prometheus_client import Counter, Histogram, Gauge
emails_sent_total = Counter(
'emails_sent_total',
'Total emails accepted by provider',
['provider', 'priority', 'campaign_type']
)
provider_send_seconds = Histogram(
'provider_send_seconds',
'Latency of provider send API',
['provider'],
buckets=(0.05, 0.1, 0.25, 0.5, 1, 2, 5, 10)
)
queue_depth = Gauge(
'queue_depth',
'Current queue depth',
['queue_name']
)
Note when using labels
- Only add labels with controllable cardinality.
- Do not attach a good email address
message_idGo to metric labels. - With campaigns, it is usually a good idea to aggregate accordingly
campaign_typeor top-N campaigns, not all campaigns at the same time.
4. What should the dashboard display?
Campaign management dashboard
- Current send rate according to provider.
- Delivery rate and bounce rate by domain.
- ETA completes the campaign.
- Queue backlog and retry backlog.
- Complaint rate in 5 minutes.
Dashboard for on-call engineers
- Top error classes in the last 15 minutes.
- Circuit breaker state of each provider.
- Worker pod restarts.
- DB latency, Redis latency.
- Webhook ingestion lag.
Dashboard for deliverability owner
- Open/click trends by domain.
- Hard bounce/complaint by segment.
- IP warming progress.
- Domain reputation indicators.
5. Alerting: few but correct warnings
Example alert rules
groups:
- name: notification-alerts
rules:
- alert: NotificationDLQSpike
expr: increase(dlq_messages_total[5m]) > 200
for: 10m
labels:
severity: page
annotations:
summary: "DLQ spike detected"
- alert: ProviderThrottleRateHigh
expr: rate(retry_attempt_total{error_class="provider_throttled"}[5m]) > 20
for: 5m
labels:
severity: warning
annotations:
summary: "Provider throttling rate is high"
- alert: CriticalQueueBacklog
expr: queue_depth{queue_name="critical-email"} > 1000
for: 2m
labels:
severity: page
annotations:
summary: "Critical email backlog exceeded threshold"
Alerting principles
- Page only when there is SLA risk or data loss.
- Warning when you need to observe but do not need to wake up on-call.
- Alerts must lead to a clear action, not just a general notification.
6. Distributed tracing with OpenTelemetry
An email goes through many hops: API -> queue -> worker -> provider -> webhook -> analytics. Without trace, it is difficult to connect the story of a message when there is an incident.
Propagate correlation identifiers
from opentelemetry import trace
tracer = trace.get_tracer(__name__)
def enqueue_email(job):
with tracer.start_as_current_span("enqueue_email") as span:
span.set_attribute("campaign.id", job.campaign_id)
span.set_attribute("message.id", job.message_id)
span.set_attribute("priority", job.priority)
publish_to_queue(job)
Fields should propagate
trace_idcampaign_idmessage_idproviderrecipient_domain
Trace does not replace metrics, but is extremely useful for analyzing a specific error or latency outlier.
7. SLO/SLA for notification system
Practical SLO example
| Traffic type | SLO |
|---|---|
| OTP / password reset | 99% sent to provider in < 15 seconds |
| Order confirmation | 99% in < 2 minutes |
| Marketing campaigns | 95% completed within committed ETA |
Error budget mindset
If your marketing workload is consuming too much of your system's error budget, you must de-prioritize or lengthen your campaign runtime. SLO helps teams make decisions using data instead of emotional arguments.
8. Runbook for common problems
Incident: queue backlog increased sharply
- Check provider throttling rate.
- Check if the autoscaler scales up.
- Check if the Redis limiter is locked too tight.
- If it is a marketing campaign, consider reducing the send rate or pausing.
Incident: complaint rate increased abnormally
- Determine which campaign is causing the spike.
- Pause that campaign first.
- Check the segment and email content.
- Reduce domain-wide throughput if the effect is widespread.
Incident: delivery decreased but sent is still high
- Check if webhooks are missing or slow.
- Check mailbox provider-specific issues.
- Comparison by domain: Is Gmail or Outlook affected separately?
- Check reputation dashboards.
Summary
Good monitoring helps you see the notification system from a real operational perspective: throughput, quality, SLA and risk. The right metrics will shorten investigation time, reduce false alarms, and allow campaigns to be scaled with much greater confidence.
Next article: We close the series with a production case study, deploying a system to send 10 million end-to-end emails on a real-world infrastructure.