Introduction
At scale, errors are no longer exceptions but the default state of the system. SMTP timeout, provider 429, slow DNS lookup, webhook not returning, database lock contention, payload data error. To be production-ready, you must design error handling as a central part of the architecture.
1. Classify errors before talking about retry
Retry is only effective when the error is temporary. If the error is permanent and still retries, you are creating a retry storm.
Error taxonomy
| Error group | Example | Is there a retry? | Action |
|---|---|---|---|
| Transient infrastructure | network timeout, DNS timeout | Yes | exponential backoff |
| Provider throttling | HTTP 429, SMTP 421 | Yes | deceleration + retry |
| Temporary mailbox issues | mailbox full, greylisting | Limited | slow retry |
| Permanent recipient failure | user unknown, invalid domain | No | mark bounced |
| Content/policy failure | blocked content, policy violation | No | move to manual review |
| Internal bugs | null field, bad template render | Not immediately | isolate, DLQ |
Important rule
- Retry according to type of error, not retry based on emotion.
- Save the original error code from the provider to provide enough classification data.
- Define clear mapping from provider error to internal error class.
2. Reasonable Retry policy
Policy matrix example
| Error class | Max retries | Delay policy |
|---|---|---|
| network_timeout | 5 | 5s, 15s, 30s, 60s, 120s |
| provider_throttled | 8 | 10s + jitter, associated with adaptive throttling |
| soft_bounce | 3 | 15m, 1h, 6h |
| rendering_failure | 0 | go straight DLQ |
| invalid_recipient | 0 | mark permanent failure |
Exponential backoff with jitter
import random
def next_retry_delay(base_seconds: int, attempt: int, max_seconds: int = 3600) -> int:
delay = min(base_seconds * (2 ** (attempt - 1)), max_seconds)
jitter = random.randint(0, max(1, delay // 4))
return delay + jitter
for attempt in range(1, 6):
print(attempt, next_retry_delay(5, attempt))
Jitter helps workers not retry an error at the same time and create additional spikes on the provider.
3. Retry queue and scheduled reprocessing
Why not immediately push back to the main queue?
If the job fails and immediately returns to the main queue, the worker can immediately reload, leading to a useless hot loop. Retry requires intentional timeout.
Popular model
send-queue
│ success
├──────────────▶ sent-status
│
│ transient failure
▼
retry-scheduler
│
├── retry in 5s
├── retry in 1m
└── retry in 15m
│
▼
send-queue
│
└── after max retry -> DLQ
Redis Sorted Set for delayed retry
def schedule_retry(redis_client, message_id: str, payload: dict, retry_at_epoch: int):
redis_client.zadd(
'email:retry:zset',
{json.dumps({'message_id': message_id, 'payload': payload}): retry_at_epoch}
)
A separate scheduler will poll the due items and republish them to the main queue.
4. How is Dead Letter Queue designed?
DLQ is not a landfill. It is a quarantine area for investigation, repair and controlled reprocessing.
When does the message enter DLQ?
- Too many retries.
- Payload is broken or violates the schema.
- Template rendering error repeats.
- Provider returns policy/permanent failure error but still needs audit.
- Unknown error class does not have a safe mapping.
Content to keep in DLQ message
{
"message_id": "msg_019e7a10_dead1",
"campaign_id": "camp_flash_sale_april",
"recipient": "[email protected]",
"error_class": "rendering_failure",
"error_code": "TEMPLATE_VARIABLE_MISSING",
"error_message": "variable first_name is required",
"attempt_count": 3,
"last_failed_at": "2026-04-01T14:00:00Z",
"original_payload": {
"template_id": "flash_sale_v2",
"variables": {"discount_code": "FLASH30"}
}
}
DLQ operating principles
- DLQ must be searchable by campaign, provider, error class.
- Has dashboard top error types.
- Reprocess is only allowed through tool/flow control.
- Do not automatically re-drive the entire DLQ if you do not understand the root cause.
5. Poison message and isolation
A poison message is a message that causes a worker to crash or fail to repeat indefinitely. If not isolated, it will destroy the entire consumer group.
How to handle it
- Very low attempt limit for unknown/internal errors.
- Wrap the parse/render/send step in a clear boundary.
- Record
error_class = poison_messageif the worker fails multiple times on the same payload. - Switch to DLQ and alert the engineer on duty.
Defensive validation before sending
def validate_job(job):
required_fields = ['message_id', 'recipient', 'template_id', 'campaign_id']
missing = [field for field in required_fields if not job.get(field)]
if missing:
raise ValueError(f"Missing required fields: {missing}")
if '@' not in job['recipient']['email']:
raise ValueError('Invalid recipient email')
return True
Early validation helps catch errors before they get deeper into the pipeline.
6. Circuit breaker for external providers
When the provider is experiencing widespread errors, constant retries only make the situation worse.
State machine is simple
| State | Meaning | Action |
|---|---|---|
| closed | normal provider | send normally |
| open | provider error exceeds threshold | stop traffic, switch to failover |
| half-open | try again with less traffic | recovery review |
Pseudocode
class CircuitBreaker:
def __init__(self, failure_threshold=10, recovery_seconds=60):
self.failure_threshold = failure_threshold
self.recovery_seconds = recovery_seconds
self.failures = 0
self.state = 'closed'
def on_success(self):
self.failures = 0
self.state = 'closed'
def on_failure(self):
self.failures += 1
if self.failures >= self.failure_threshold:
self.state = 'open'
The circuit breaker must be attached to the provider failover strategy. If the breaker is not opened without an alternative path, the queue will still just stand still.
7. Compensation logic and consistent state
An email may have been received by the provider, but the worker failed before updating the DB. Without compensation, you won't know whether an email has been sent or not.
How to handle it
- Mount
message_idinternal to the metadata sent by the provider. - Receive webhook or callback to reconcile state.
- Have periodic job reconciliation: compare outbound log with provider events.
- For uncertain status, check
unknowninstead of assuming failure.
Recommended internal status
pending -> processing -> sent_to_provider -> delivered
│
├-> deferred
├-> bounced_permanent
├-> bounced_transient
└-> failed_internal
8. Alerting and minimal runbook
Warning should be there
- The rate of messages entering DLQ increased dramatically.
- An error class accounts for over 20% in 5 minutes.
- Circuit breaker opens with main provider.
- Retry queue backlog increases but throughput decreases.
- New unknown error code appears.
Short runbook for on-call
- Identify the error in the new provider, data, or deployed code.
- Check if you need to pause campaign or failover provider.
- View the top error class and some sample payloads in DLQ.
- Only re-drive after verifying that the root cause has been resolved.
Summary
Retry and DLQ are not aftermarket accessories. They are a mandatory mechanism for the system to tolerate errors without creating chaos. Good design must know how to classify errors, retry strategically, isolate poison messages, and leave enough traces for investigation.
Next article: We move from the ability to send to the ability to reach the inbox, that is, deliverability with SPF, DKIM, DMARC and reputation management.