Chuyển đến nội dung chính

第 30 課:架構決策記錄與生產清單

架構決策記錄(ADR):格式、何時編寫、實例。生產準備清單。系統設計面試綜合框架。系統架構師的職業道路。資源和後續步驟。

🏗️ 建築 — 第 30 課 第 30 課:架構決策記錄與 生產清單

系統架構:從零到英雄

第 8 部分:生產就緒架構

亞洲開發網

簡介

最後一課 - 將學到的所有知識綜合到實用工具中:架構決策記錄用於記錄決策,生產清單用於確保系統準備就緒,以及系統設計框架用於訪談。


1. 架構決策記錄(ADR)

1.1 為什麼需要 ADR?

6 tháng sau:
  Developer mới: "Tại sao chúng ta dùng MongoDB thay vì PostgreSQL?"
  Team: "Hmm... không ai nhớ."

  Với ADR:
  → Đọc ADR-005: "Chọn MongoDB vì product catalog cần flexible schema,
     variants có attributes khác nhau. PostgreSQL JSONB cũng được xem xét
     nhưng MongoDB có better query performance cho nested documents."

ADR = Document GHI LẠI quyết định kiến trúc quan trọng
  - Context: Tại sao phải quyết định?
  - Decision: Quyết định gì?
  - Consequences: Hậu quả là gì?

1.2 ADR 模板

# ADR-001: Chọn Message Queue

## Status
Accepted (2024-01-15)

## Context
Hệ thống cần async processing cho: email notifications,
order processing, analytics events. Hiện tại tất cả
synchronous, gây latency 2-3 giây cho checkout.

## Decision Drivers
- Throughput: 10K messages/s peak
- Ordering: Per-partition ordering needed
- Retention: Need message replay (audit)
- Team expertise: Team có kinh nghiệm Kafka

## Options Considered
1. RabbitMQ: Mature, easy setup, good routing
2. Apache Kafka: High throughput, retention, replay
3. AWS SQS: Managed, simple, no operations

## Decision
Chọn Apache Kafka.

## Rationale
- Throughput requirement (10K/s) → Kafka excels
- Message retention + replay → Audit compliance
- Team already has Kafka experience
- Event-driven architecture direction → Kafka fits

## Consequences
### Positive
- Message replay for debugging/audit
- High throughput headroom
- Foundation for event-driven architecture

### Negative
- Operational complexity (ZooKeeper/KRaft cluster)
- Higher learning curve for new team members
- Need monitoring setup (Kafka lag, consumer groups)

### Risks
- Kafka cluster management overhead → Mitigate: Use managed Kafka (MSK/Confluent)

## References
- Bài 13: Message Queues & Task Queues
- RFC-2024-003: Async Processing Architecture

1.3 什麼時候寫ADR?

✅ Viết ADR khi:
  - Chọn technology/framework mới
  - Thay đổi architecture pattern
  - Quyết định ảnh hưởng nhiều teams
  - Trade-off rõ ràng (security vs performance)
  - Quyết định khó đảo ngược

❌ KHÔNG cần ADR khi:
  - Naming convention (dùng coding standards)
  - Library minor version
  - Implementation details (code review đủ)
  - Temporary solutions (dùng TODO/RFC)

2. 生產準備清單

2.1 基礎設施

□ Auto-scaling configured (min/max/target)
□ Multi-AZ deployment
□ Load balancer health checks
□ Database replicas + failover tested
□ Backup strategy verified (test restore!)
□ CDN for static assets
□ DNS with TTL configured
□ SSL/TLS certificates (auto-renew)
□ Secret management (Vault/KMS, not env vars)
□ Infrastructure as Code (Terraform/Pulumi)

2.2 應用

□ Health check endpoints (/healthz, /readyz)
□ Graceful shutdown (drain connections)
□ Connection pooling (database, HTTP)
□ Timeouts on ALL external calls
□ Circuit breakers for downstream services
□ Retry with exponential backoff
□ Rate limiting configured
□ Input validation (all endpoints)
□ Error handling (no stack traces in production)
□ Feature flags for risky deployments

2.3 可觀察性

□ Metrics: RED (Rate, Error, Duration) per service
□ Logging: Structured, centralized (ELK/Loki)
□ Tracing: Distributed tracing (OpenTelemetry)
□ Dashboards: Service overview, business metrics
□ Alerts: P1/P2 alerts with runbooks
□ On-call rotation configured
□ Status page for external communication
□ SLO defined and monitored

2.4 安全

□ Authentication + Authorization
□ HTTPS everywhere, HSTS headers
□ WAF configured
□ Dependency vulnerability scanning
□ Container image scanning
□ Secrets rotation policy
□ Audit logging for sensitive operations
□ Penetration testing done
□ Data encryption at rest and in transit
□ CORS, CSP headers configured

2.5 可靠性

□ Chaos experiments run in staging
□ Disaster Recovery plan documented
□ DR failover tested
□ Runbooks for common incidents
□ Post-incident review process
□ Capacity planning for next 6 months
□ Load testing done (normal + peak)
□ Database migration strategy (zero-downtime)
□ Rollback plan for every deployment
□ Data consistency checks automated

2.6 持續整合/持續交付

□ Automated tests (unit, integration, e2e)
□ Code coverage > 80%
□ Linting + formatting enforced
□ Container build pipeline
□ Staging environment mirrors production
□ Blue-green or canary deployment
□ Automated rollback on failure
□ Database migration in pipeline
□ Security scanning in pipeline
□ Performance regression tests

3.系統設計面試框架

3.1 4 步驟框架(45 分鐘)

Step 1: Requirements Clarification (5 min)
  - Functional requirements (features)
  - Non-functional requirements (QPS, latency, availability)
  - Scale estimation (users, data, bandwidth)
  - Constraints (existing systems, compliance)

Step 2: High-Level Design (10 min)
  - Draw main components
  - Show data flow
  - Identify key services
  - API design (endpoints, request/response)

Step 3: Deep Dive (20 min)
  - Database schema + choice justification
  - Scaling strategy (sharding, caching)
  - Algorithm design (matching, ranking)
  - Handle edge cases (race conditions, failures)

Step 4: Wrap Up (10 min)
  - Bottleneck identification
  - Scaling discussion (10x, 100x)
  - Monitoring + alerting approach
  - Trade-offs recap

3.2 估算備忘單

Numbers Everyone Should Know:
  L1 cache reference:     0.5 ns
  L2 cache reference:     7 ns
  Memory reference:       100 ns
  SSD random read:        150 μs
  HDD random read:        10 ms
  Network round trip:     500 μs (same DC)
  Network round trip:     150 ms (cross-continent)

Storage:
  1 char = 1 byte (ASCII) / 2-4 bytes (UTF-8)
  1 UUID = 36 bytes (string) / 16 bytes (binary)
  1 timestamp = 8 bytes
  Average tweet = 140 bytes
  Average URL = 100 bytes
  Average image = 300KB
  Average video (1 min) = 50MB

Scale:
  1M seconds ≈ 12 days
  1B seconds ≈ 32 years
  QPS → 86,400 = requests per day (÷86,400 = QPS)
  2^10 = 1K, 2^20 = 1M, 2^30 = 1G, 2^40 = 1T

4. 回顧:30 課

Foundation (1-4):
  System Design overview, performance, scalability,
  CAP theorem, networking

Infrastructure (5-8):
  Load balancer, CDN, caching, API gateway

Database (9-12):
  SQL vs NoSQL, replication, sharding, storage patterns

Async (13-15):
  Message queues, event-driven, stream processing

Architecture (16-19):
  Microservices, service communication, DDD, serverless

Reliability (20-23):
  HA, DR, security, observability

Case Studies (24-29):
  URL shortener, chat, news feed, video streaming,
  ride-sharing, e-commerce

Production (30):
  ADR, production checklist, interview framework

5.學習資源

Books:
  📚 "Designing Data-Intensive Applications" - Martin Kleppmann
  📚 "System Design Interview" - Alex Xu (Vol 1 & 2)
  📚 "Building Microservices" - Sam Newman
  📚 "Domain-Driven Design" - Eric Evans
  📚 "Site Reliability Engineering" - Google

Online:
  🌐 system-design-primer (GitHub)
  🌐 ByteByteGo (Alex Xu)
  🌐 roadmap.sh/system-design
  🌐 highscalability.com

Practice:
  💻 Design 1 system per week
  💻 Write ADRs for your current projects
  💻 Read engineering blogs: Netflix, Uber, Shopify
  💻 Contribute to open-source infra projects

6. 職業道路

Junior Developer
  → Understand single-server apps
  → Learn database fundamentals

Mid-Level Developer
  → Design multi-server systems
  → Implement caching, queues
  → Operational experience

Senior Developer
  → Lead architecture decisions
  → Write ADRs
  → Mentor on system design

Staff/Principal Engineer
  → Cross-team architecture
  → Define tech strategy
  → Influence org-wide patterns

System Architect
  → Enterprise architecture
  → Vendor evaluation
  → Multi-year technology roadmap

結論

您已經完成了「系統架構:從零到英雄」—30 節課,80 小時的學習。系統架構不只是理論,更是理論。請套用到實際項目。為您的下一個項目編寫 ADR,查看生產清單,並每天繼續學習。

「最好的架構是能夠簡單地解決問題的架構。」