Chuyển đến nội dung chính

Lesson 14: Monitoring, Metrics & Alerting

Key metrics for notification platform, Prometheus and Grafana dashboard, distributed tracing with OpenTelemetry, queue depth monitoring, worker health checks, SLA tracking and runbooks for common incidents.

🏗️ Architecture — Lesson 14 Lesson 14: Monitoring, Metrics & Alerting

Design a Notification System to send millions of Emails

Part 5: Deliverability, Monitoring & Production

xdev.asia

Introduction

Without metrics, a 10 million email campaign is just a leap of faith into the dark. The notification production system must be able to immediately answer the questions: how fast is it sending, where is it stuck, which provider is failing, which campaign is about to break its SLA.


1. Observe the system layer by layer

Four main layers of metrics

ClassFor example metric
Businessemails sent, delivered, complaints, campaign ETA
Applicationrender latency, send latency, retry count
Queueconsumer lag, queue depth, DLQ size
InfrastructureCPU, memory, network, pod restarts

Anti-patterns should be avoided

  • Only look at CPU/RAM, not queue lag.
  • Just look sent_count without looking at delivery/bounce/complaint.
  • Do not attach metrics provider, campaign_type, recipient_domain.

2. Core set of metrics for email platform

Business metrics

MetricsMeaning
emails_requested_totaltotal number of emails requested to be sent
emails_sent_totalnumber of emails accepted by provider
emails_delivered_totaldelivery confirmed email number
emails_bounced_totalbounce by type
emails_complained_totalcomplaint spam
campaign_eta_secondsCampaign completion ETA

Application metrics

MetricsLabels should have
template_render_secondstemplate_id, locale
provider_send_secondsprovider, priority
retry_attempt_totalerror_class, provider
throttle_decision_totallimit_type, result
worker_batch_duration_secondsworker_group

Queue metrics

MetricsMeaning
queue_depthtotal pending jobs
consumer_lagbacklog level compared to producer
retry_queue_depthnumber of jobs waiting for retry
dlq_messages_totalquantity goes into DLQ

3. Prometheus instrumentation

Example metric definitions

from prometheus_client import Counter, Histogram, Gauge

emails_sent_total = Counter(
    'emails_sent_total',
    'Total emails accepted by provider',
    ['provider', 'priority', 'campaign_type']
)

provider_send_seconds = Histogram(
    'provider_send_seconds',
    'Latency of provider send API',
    ['provider'],
    buckets=(0.05, 0.1, 0.25, 0.5, 1, 2, 5, 10)
)

queue_depth = Gauge(
    'queue_depth',
    'Current queue depth',
    ['queue_name']
)

Note when using labels

  • Only add labels with controllable cardinality.
  • Do not attach a good email address message_id Go to metric labels.
  • With campaigns, it is usually a good idea to aggregate accordingly campaign_type or top-N campaigns, not all campaigns at the same time.

4. What should the dashboard display?

Campaign management dashboard

  • Current send rate according to provider.
  • Delivery rate and bounce rate by domain.
  • ETA completes the campaign.
  • Queue backlog and retry backlog.
  • Complaint rate in 5 minutes.

Dashboard for on-call engineers

  • Top error classes in the last 15 minutes.
  • Circuit breaker state of each provider.
  • Worker pod restarts.
  • DB latency, Redis latency.
  • Webhook ingestion lag.

Dashboard for deliverability owner

  • Open/click trends by domain.
  • Hard bounce/complaint by segment.
  • IP warming progress.
  • Domain reputation indicators.

5. Alerting: few but correct warnings

Example alert rules

groups:
  - name: notification-alerts
    rules:
      - alert: NotificationDLQSpike
        expr: increase(dlq_messages_total[5m]) > 200
        for: 10m
        labels:
          severity: page
        annotations:
          summary: "DLQ spike detected"

      - alert: ProviderThrottleRateHigh
        expr: rate(retry_attempt_total{error_class="provider_throttled"}[5m]) > 20
        for: 5m
        labels:
          severity: warning
        annotations:
          summary: "Provider throttling rate is high"

      - alert: CriticalQueueBacklog
        expr: queue_depth{queue_name="critical-email"} > 1000
        for: 2m
        labels:
          severity: page
        annotations:
          summary: "Critical email backlog exceeded threshold"

Alerting principles

  • Page only when there is SLA risk or data loss.
  • Warning when you need to observe but do not need to wake up on-call.
  • Alerts must lead to a clear action, not just a general notification.

6. Distributed tracing with OpenTelemetry

An email goes through many hops: API -> queue -> worker -> provider -> webhook -> analytics. Without trace, it is difficult to connect the story of a message when there is an incident.

Propagate correlation identifiers

from opentelemetry import trace

tracer = trace.get_tracer(__name__)

def enqueue_email(job):
    with tracer.start_as_current_span("enqueue_email") as span:
        span.set_attribute("campaign.id", job.campaign_id)
        span.set_attribute("message.id", job.message_id)
        span.set_attribute("priority", job.priority)
        publish_to_queue(job)

Fields should propagate

  • trace_id
  • campaign_id
  • message_id
  • provider
  • recipient_domain

Trace does not replace metrics, but is extremely useful for analyzing a specific error or latency outlier.


7. SLO/SLA for notification system

Practical SLO example

Traffic typeSLO
OTP / password reset99% sent to provider in < 15 seconds
Order confirmation99% in < 2 minutes
Marketing campaigns95% completed within committed ETA

Error budget mindset

If your marketing workload is consuming too much of your system's error budget, you must de-prioritize or lengthen your campaign runtime. SLO helps teams make decisions using data instead of emotional arguments.


8. Runbook for common problems

Incident: queue backlog increased sharply

  1. Check provider throttling rate.
  2. Check if the autoscaler scales up.
  3. Check if the Redis limiter is locked too tight.
  4. If it is a marketing campaign, consider reducing the send rate or pausing.

Incident: complaint rate increased abnormally

  1. Determine which campaign is causing the spike.
  2. Pause that campaign first.
  3. Check the segment and email content.
  4. Reduce domain-wide throughput if the effect is widespread.

Incident: delivery decreased but sent is still high

  1. Check if webhooks are missing or slow.
  2. Check mailbox provider-specific issues.
  3. Comparison by domain: Is Gmail or Outlook affected separately?
  4. Check reputation dashboards.

Summary

Good monitoring helps you see the notification system from a real operational perspective: throughput, quality, SLA and risk. The right metrics will shorten investigation time, reduce false alarms, and allow campaigns to be scaled with much greater confidence.

Next article: We close the series with a production case study, deploying a system to send 10 million end-to-end emails on a real-world infrastructure.