Chuyển đến nội dung chính

Lesson 1: What is System Design? - Overview and Roadmap

Introducing System Design, why system design is needed, how to approach a system design problem (requirements → high-level design → deep dive → bottlenecks). Compare Monolith vs Distributed Systems. Learning roadmap and necessary resources.

🏗️ Architecture — Lesson 1 Lesson 1: What is System Design? - Overview and Roadmap

System Architecture: From Zero to Hero

Part 1: System Design Foundation

xdev.asia

Introduction

You can write a simple CRUD application in a few hours. But when that application needs to serve millions of users, handle thousands of requests/second, and ensure 99.99% uptime — that's when you need System Design.

System Design is not just knowledge for interviews. It's a core skill that helps you:

  • Build scalable systems
  • Make the right architectural decisions
  • Avoid costly mistakes (refactor the entire system)
  • Communicate effectively with the team about technical decisions

1. What is System Design?

1.1 Definition

System Design is the process of determining the architecture, components, modules, interfaces and data for a software system to satisfy specific requirements.

System Design = Architecture + Components + Data Flow + Trade-offs

1.2 Why is System Design important?

PhaseNo System DesignYes System Design
PrototypeFast, simpleHave a clear plan
100 usersRuns wellRuns well
10K usersSlow startStill stable
1M usersSystem crash, had to rewriteScale according to plan
CostRewrite = 10x initial costIncremental improvements

1.3 System Design vs Coding

Coding:          "Làm sao để implement feature X?"
System Design:   "Làm sao để feature X hoạt động với 10M users,
                  99.99% uptime, <100ms latency?"

2. How to approach the System Design problem

2.1 Framework 4 steps

When faced with any system design problem, use the following framework:

┌─────────────────────────────────────────────────────┐
│  Step 1: Requirements & Constraints                  │
│  ┌─────────────────────────────────────────────┐    │
│  │ Functional Requirements (FR)                 │    │
│  │ Non-functional Requirements (NFR)            │    │
│  │ Constraints & Assumptions                    │    │
│  └─────────────────────────────────────────────┘    │
│                      ▼                               │
│  Step 2: High-Level Design                           │
│  ┌─────────────────────────────────────────────┐    │
│  │ Main components & connections                │    │
│  │ Data flow diagrams                           │    │
│  │ API design (endpoints)                       │    │
│  └─────────────────────────────────────────────┘    │
│                      ▼                               │
│  Step 3: Deep Dive into Core Components              │
│  ┌─────────────────────────────────────────────┐    │
│  │ Database schema                              │    │
│  │ Algorithm choices                            │    │
│  │ Data structures                              │    │
│  └─────────────────────────────────────────────┘    │
│                      ▼                               │
│  Step 4: Identify & Resolve Bottlenecks              │
│  ┌─────────────────────────────────────────────┐    │
│  │ Single points of failure                     │    │
│  │ Scaling strategies                           │    │
│  │ Monitoring & alerting                        │    │
│  └─────────────────────────────────────────────┘    │
└─────────────────────────────────────────────────────┘

2.2 Example: Designing a news reading system

Step 1 - Requirements:

  • FR: User views article list, reads details, searches, comments
  • NFR: 10M DAYS, <200ms latency, 99.9% availability
  • Constraints: Read-heavy (100:1 read/write ratio)

Step 2 - High-Level Design:

Users → CDN → Load Balancer → Web Servers → Cache → Database
                    │
                    └→ Search Service (Elasticsearch)

Step 3 - Deep Dive:

  • Database: PostgreSQL cho articles, Redis cho cache
  • Search: Elasticsearch with full-text index
  • CDN: Cache static assets + rendered HTML

Step 4 - Bottlenecks:

  • Database read bottleneck → Add read replicas
  • Hot articles → Aggressive caching with TTL
  • Search latency → Elasticsearch cluster scaling

3. Monolith vs Distributed Systems

3.1 Monolithic Architecture

┌─────────────────────────────────────┐
│         Monolithic Application       │
│  ┌─────┬─────┬─────┬─────┬──────┐  │
│  │ UI  │User │Order│Pay  │Search│  │
│  │Layer│ Svc │ Svc │ Svc │ Svc  │  │
│  └─────┴─────┴─────┴─────┴──────┘  │
│  ┌─────────────────────────────┐    │
│  │      Shared Database         │    │
│  └─────────────────────────────┘    │
└─────────────────────────────────────┘

Advantage:

  • Simple to initially develop
  • Easy to deploy (1 artifact)
  • Easy to debug (single process)
  • No network latency between components

Disadvantages:

  • Difficult to scale each individual part
  • A small error can crash the entire system
  • Deploy is slow when the codebase is large
  • Technology lock-in (1 language/framework)

3.2 Distributed Systems

┌────────┐  ┌────────┐  ┌────────┐  ┌────────┐
│ User   │  │ Order  │  │Payment │  │ Search │
│Service │  │Service │  │Service │  │Service │
│  DB    │  │  DB    │  │  DB    │  │  DB    │
└───┬────┘  └───┬────┘  └───┬────┘  └───┬────┘
    │           │           │           │
    └───────────┴─────┬─────┴───────────┘
                      │
              Message Queue / API Gateway

Advantage:

  • Scale each service independently
  • Fault isolation (1 service failure ≠ whole system failure)
  • Team autonomy (each team owns 1 service)
  • Technology diversity

Disadvantages:

  • Much more complicated
  • Network latency between services
  • Data consistency challenges
  • Operational overhead (monitoring, debugging)

3.3 When to choose what?

CriteriaMonolithDistributed
Team size< 10 developers> 10 developers
Traffic< 10K RPS> 10K RPS
PhaseMVP, Startup earlyGrowth, Scale
ComplexityModerateHigh
Deploy frequencyWeekly/MonthlyDaily/Hourly

Advice: Most systems should start with Monolith, then branch out to distributed as needed. Don't over-engineer from the beginning!


4. Core concepts to grasp

4.1 System Design Map

                    System Design
                         │
    ┌────────────────────┼────────────────────┐
    │                    │                    │
Fundamentals        Components           Patterns
    │                    │                    │
├─ Scalability      ├─ Load Balancer    ├─ Microservices
├─ Availability     ├─ CDN              ├─ Event-Driven
├─ Consistency      ├─ Cache            ├─ CQRS
├─ Latency          ├─ Database         ├─ Saga
├─ Throughput       ├─ Message Queue    ├─ Circuit Breaker
├─ CAP Theorem      ├─ API Gateway      ├─ DDD
└─ Networking       ├─ Reverse Proxy    └─ Serverless
                    └─ Search Engine

4.2 Learning Roadmap

Tháng 1-2: Fundamentals
  ├─ Scalability, Availability, Consistency
  ├─ CAP Theorem
  └─ Networking basics

Tháng 3-4: Infrastructure Components
  ├─ Load Balancer, CDN, Cache
  ├─ Database (SQL, NoSQL, Sharding)
  └─ Message Queues

Tháng 5-6: Architectural Patterns
  ├─ Microservices, Event-Driven
  ├─ CQRS, Saga, DDD
  └─ Serverless

Tháng 7-8: Case Studies & Practice
  ├─ Design URL Shortener
  ├─ Design Chat System
  ├─ Design News Feed
  └─ Design Video Streaming

5. Back-of-the-Envelope Calculations

An important skill in System Design is rapid estimation:

5.1 Powers of Two

PowerExact ValueApproxBytes
101,0241 Thousand1 KB
201,048,5761 Million1MB
301,073,741,8241 Billion1GB
401,099,511,627,7761 Trillion1 TB

5.2 Latency Numbers Every Programmer Should Know

L1 cache reference:                    0.5 ns
L2 cache reference:                      7 ns
Main memory reference:                 100 ns
SSD random read:                   150,000 ns  =  150 μs
HDD seek:                      10,000,000 ns  =   10 ms
Send 1 MB over 1 Gbps network: 10,000,000 ns  =   10 ms
Read 1 MB from SSD:             1,000,000 ns  =    1 ms
Read 1 MB from HDD:            30,000,000 ns  =   30 ms
Roundtrip same datacenter:        500,000 ns  =  500 μs
Roundtrip CA → Netherlands:   150,000,000 ns  =  150 ms

5.3 Example: Storage estimates for Twitter

Giả sử:
- 500M users, 200M DAU
- Mỗi user tweet 2 lần/ngày
- Mỗi tweet: 140 chars * 2 bytes = 280 bytes
- 10% tweets có media (ảnh 200KB trung bình)

Tweets/ngày: 200M * 2 = 400M tweets
Text storage/ngày: 400M * 280B = 112 GB/ngày
Media storage/ngày: 40M * 200KB = 8 TB/ngày

Storage/năm: (112GB + 8TB) * 365 ≈ 3 PB/năm

6. Summary

TopicsKey Takeaway
System DesignDesign systems that satisfy requirements on a large scale
FrameworksRequirements → High-Level → Deep Dive → Bottlenecks
MonolithStart with monolith, separate as needed
DistributedMore complicated but allows scaling
EstimationAlways estimate before designing

Exercises

  1. Estimation Practice: Estimated storage needed for YouTube in 1 year (500M DAU, 5M videos uploaded/day, average 50MB/video after transcoding)

  2. Monolith vs Distributed: You are building a hotel booking application for the Vietnamese market (5M users). Will you choose Monolith or Distributed? Explain why.

  3. System Design Framework: Apply a 4-step framework to sketch the design for the restaurant reservation management system (50K restaurants, 1M users).