Chuyển đến nội dung chính

レッスン 30: アーキテクチャの決定記録と制作チェックリスト

アーキテクチャ決定記録 (ADR): 形式、いつ作成するか、実際の例。本番準備チェックリスト。システム設計面接の包括的なフレームワーク。システムアーキテクトのキャリアパス。リソースと次のステップ。

🏗️ アーキテクチャ — レッスン 30 レッスン 30: アーキテクチャの決定記録と 生産チェックリスト

システムアーキテクチャ: ゼロからヒーローへ

パート 8: 本番環境に対応したアーキテクチャ

xdev.asia

はじめに

最後のレッスン — 学んだことすべてを実用的なツールに統合します。決定事項を文書化するための アーキテクチャ決定記録、システムの準備を確実にするための 運用チェックリスト、およびインタビューのための システム設計フレームワーク。


1. アーキテクチャ決定記録 (ADR)

1.1 ADR はなぜ必要ですか?

6 tháng sau:
  Developer mới: "Tại sao chúng ta dùng MongoDB thay vì PostgreSQL?"
  Team: "Hmm... không ai nhớ."

  Với ADR:
  → Đọc ADR-005: "Chọn MongoDB vì product catalog cần flexible schema,
     variants có attributes khác nhau. PostgreSQL JSONB cũng được xem xét
     nhưng MongoDB có better query performance cho nested documents."

ADR = Document GHI LẠI quyết định kiến trúc quan trọng
  - Context: Tại sao phải quyết định?
  - Decision: Quyết định gì?
  - Consequences: Hậu quả là gì?

1.2 ADR テンプレート

# ADR-001: Chọn Message Queue

## Status
Accepted (2024-01-15)

## Context
Hệ thống cần async processing cho: email notifications,
order processing, analytics events. Hiện tại tất cả
synchronous, gây latency 2-3 giây cho checkout.

## Decision Drivers
- Throughput: 10K messages/s peak
- Ordering: Per-partition ordering needed
- Retention: Need message replay (audit)
- Team expertise: Team có kinh nghiệm Kafka

## Options Considered
1. RabbitMQ: Mature, easy setup, good routing
2. Apache Kafka: High throughput, retention, replay
3. AWS SQS: Managed, simple, no operations

## Decision
Chọn Apache Kafka.

## Rationale
- Throughput requirement (10K/s) → Kafka excels
- Message retention + replay → Audit compliance
- Team already has Kafka experience
- Event-driven architecture direction → Kafka fits

## Consequences
### Positive
- Message replay for debugging/audit
- High throughput headroom
- Foundation for event-driven architecture

### Negative
- Operational complexity (ZooKeeper/KRaft cluster)
- Higher learning curve for new team members
- Need monitoring setup (Kafka lag, consumer groups)

### Risks
- Kafka cluster management overhead → Mitigate: Use managed Kafka (MSK/Confluent)

## References
- Bài 13: Message Queues & Task Queues
- RFC-2024-003: Async Processing Architecture

1.3 ADR をいつ作成するか?

✅ Viết ADR khi:
  - Chọn technology/framework mới
  - Thay đổi architecture pattern
  - Quyết định ảnh hưởng nhiều teams
  - Trade-off rõ ràng (security vs performance)
  - Quyết định khó đảo ngược

❌ KHÔNG cần ADR khi:
  - Naming convention (dùng coding standards)
  - Library minor version
  - Implementation details (code review đủ)
  - Temporary solutions (dùng TODO/RFC)

2. 実稼働準備チェックリスト

2.1 インフラストラクチャ

□ Auto-scaling configured (min/max/target)
□ Multi-AZ deployment
□ Load balancer health checks
□ Database replicas + failover tested
□ Backup strategy verified (test restore!)
□ CDN for static assets
□ DNS with TTL configured
□ SSL/TLS certificates (auto-renew)
□ Secret management (Vault/KMS, not env vars)
□ Infrastructure as Code (Terraform/Pulumi)

2.2 アプリケーション

□ Health check endpoints (/healthz, /readyz)
□ Graceful shutdown (drain connections)
□ Connection pooling (database, HTTP)
□ Timeouts on ALL external calls
□ Circuit breakers for downstream services
□ Retry with exponential backoff
□ Rate limiting configured
□ Input validation (all endpoints)
□ Error handling (no stack traces in production)
□ Feature flags for risky deployments

2.3 可観測性

□ Metrics: RED (Rate, Error, Duration) per service
□ Logging: Structured, centralized (ELK/Loki)
□ Tracing: Distributed tracing (OpenTelemetry)
□ Dashboards: Service overview, business metrics
□ Alerts: P1/P2 alerts with runbooks
□ On-call rotation configured
□ Status page for external communication
□ SLO defined and monitored

2.4 セキュリティ

□ Authentication + Authorization
□ HTTPS everywhere, HSTS headers
□ WAF configured
□ Dependency vulnerability scanning
□ Container image scanning
□ Secrets rotation policy
□ Audit logging for sensitive operations
□ Penetration testing done
□ Data encryption at rest and in transit
□ CORS, CSP headers configured

2.5 信頼性

□ Chaos experiments run in staging
□ Disaster Recovery plan documented
□ DR failover tested
□ Runbooks for common incidents
□ Post-incident review process
□ Capacity planning for next 6 months
□ Load testing done (normal + peak)
□ Database migration strategy (zero-downtime)
□ Rollback plan for every deployment
□ Data consistency checks automated

2.6 CI/CD

□ Automated tests (unit, integration, e2e)
□ Code coverage > 80%
□ Linting + formatting enforced
□ Container build pipeline
□ Staging environment mirrors production
□ Blue-green or canary deployment
□ Automated rollback on failure
□ Database migration in pipeline
□ Security scanning in pipeline
□ Performance regression tests

3. システム設計面接の枠組み

3.1 4 ステップのフレームワーク (45 分)

Step 1: Requirements Clarification (5 min)
  - Functional requirements (features)
  - Non-functional requirements (QPS, latency, availability)
  - Scale estimation (users, data, bandwidth)
  - Constraints (existing systems, compliance)

Step 2: High-Level Design (10 min)
  - Draw main components
  - Show data flow
  - Identify key services
  - API design (endpoints, request/response)

Step 3: Deep Dive (20 min)
  - Database schema + choice justification
  - Scaling strategy (sharding, caching)
  - Algorithm design (matching, ranking)
  - Handle edge cases (race conditions, failures)

Step 4: Wrap Up (10 min)
  - Bottleneck identification
  - Scaling discussion (10x, 100x)
  - Monitoring + alerting approach
  - Trade-offs recap

3.2 見積もりチートシート

Numbers Everyone Should Know:
  L1 cache reference:     0.5 ns
  L2 cache reference:     7 ns
  Memory reference:       100 ns
  SSD random read:        150 μs
  HDD random read:        10 ms
  Network round trip:     500 μs (same DC)
  Network round trip:     150 ms (cross-continent)

Storage:
  1 char = 1 byte (ASCII) / 2-4 bytes (UTF-8)
  1 UUID = 36 bytes (string) / 16 bytes (binary)
  1 timestamp = 8 bytes
  Average tweet = 140 bytes
  Average URL = 100 bytes
  Average image = 300KB
  Average video (1 min) = 50MB

Scale:
  1M seconds ≈ 12 days
  1B seconds ≈ 32 years
  QPS → 86,400 = requests per day (÷86,400 = QPS)
  2^10 = 1K, 2^20 = 1M, 2^30 = 1G, 2^40 = 1T

4. 要約: 30 のレッスン

Foundation (1-4):
  System Design overview, performance, scalability,
  CAP theorem, networking

Infrastructure (5-8):
  Load balancer, CDN, caching, API gateway

Database (9-12):
  SQL vs NoSQL, replication, sharding, storage patterns

Async (13-15):
  Message queues, event-driven, stream processing

Architecture (16-19):
  Microservices, service communication, DDD, serverless

Reliability (20-23):
  HA, DR, security, observability

Case Studies (24-29):
  URL shortener, chat, news feed, video streaming,
  ride-sharing, e-commerce

Production (30):
  ADR, production checklist, interview framework

5. 学習リソース

Books:
  📚 "Designing Data-Intensive Applications" - Martin Kleppmann
  📚 "System Design Interview" - Alex Xu (Vol 1 & 2)
  📚 "Building Microservices" - Sam Newman
  📚 "Domain-Driven Design" - Eric Evans
  📚 "Site Reliability Engineering" - Google

Online:
  🌐 system-design-primer (GitHub)
  🌐 ByteByteGo (Alex Xu)
  🌐 roadmap.sh/system-design
  🌐 highscalability.com

Practice:
  💻 Design 1 system per week
  💻 Write ADRs for your current projects
  💻 Read engineering blogs: Netflix, Uber, Shopify
  💻 Contribute to open-source infra projects

6. キャリアパス

Junior Developer
  → Understand single-server apps
  → Learn database fundamentals

Mid-Level Developer
  → Design multi-server systems
  → Implement caching, queues
  → Operational experience

Senior Developer
  → Lead architecture decisions
  → Write ADRs
  → Mentor on system design

Staff/Principal Engineer
  → Cross-team architecture
  → Define tech strategy
  → Influence org-wide patterns

System Architect
  → Enterprise architecture
  → Vendor evaluation
  → Multi-year technology roadmap

結論

「システム アーキテクチャ: ゼロからヒーローまで」 - 30 レッスン、80 時間の学習を完了しました。システムアーキテクチャは単なる理論ではありません。ぜひ実際のプロジェクトに応用してみてください。次のプロジェクトの ADR を作成し、制作チェックリストを確認し、毎日学習を続けてください。

「最良のアーキテクチャとは、問題をシンプルに解決するものです。」