Chuyển đến nội dung chính

BÀI 11: KIẾN TRÚC DISTRIBUTED STORAGE VỚI ROOK-CEPH

Tổng quan Ceph architecture (RADOS, OSD, MON, MDS, MGR), tại sao chọn Rook-Ceph cho K8s, so sánh storage solutions, planning capacity và network cho Ceph cluster.

🔒 DevSecOps — Bài 11 BÀI 11: KIẾN TRÚC DISTRIBUTED STORAGE VỚI ROOK-CEPH

Deploy Microservices On-Premises với Kubernetes HA

Phần 3: Distributed Storage — Rook-Ceph

xdev.asia

🎯 MỤC TIÊU BÀI HỌC

Sau khi hoàn thành bài học này, bạn sẽ:

  • ✅ Hiểu kiến trúc Ceph: RADOS, OSD, MON, MDS, MGR
  • ✅ Hiểu CRUSH algorithm và data placement
  • ✅ So sánh Rook-Ceph vs Longhorn vs OpenEBS vs local-path
  • ✅ Planning hardware và network cho Ceph cluster
  • ✅ Hiểu 3 loại storage: Block (RBD), Filesystem (CephFS), Object (RGW)

PHẦN 1: CEPH ARCHITECTURE OVERVIEW

1.1. Các thành phần Ceph

graph TB
    subgraph CLIENT["🖥️ CLIENT ACCESS"]
        RBD["💿 RBD<br/>Block Storage"]
        CephFS["📁 CephFS<br/>Filesystem"]
        RGW["🌐 RADOS Gateway<br/>Object / S3-compat"]
    end

    subgraph LIB["📚 LIBRADOS API"]
        librados["Unified Storage API"]
    end

    subgraph RADOS["⚙️ RADOS LAYER"]
        MON["🔍 MON<br/>Monitor<br/>3 nodes"]
        MGR["📊 MGR<br/>Manager<br/>2 nodes"]
        MDS["📂 MDS<br/>Metadata Server<br/>CephFS only"]
        OSD["💾 OSD<br/>Object Storage<br/>Daemon"]
    end

    subgraph CRUSH["🗺️ CRUSH MAP"]
        crush_algo["Data placement algorithm — no lookup table!"]
    end

    RBD --> librados
    CephFS --> librados
    RGW --> librados
    librados --> MON
    librados --> MGR
    librados --> MDS
    librados --> OSD
    OSD --> crush_algo

    style CLIENT fill:#0f172a,stroke:#3b82f6,color:#e2e8f0
    style LIB fill:#1e3a5f,stroke:#60a5fa,color:#e2e8f0
    style RADOS fill:#1e293b,stroke:#3b82f6,color:#e2e8f0
    style CRUSH fill:#7c3aed,stroke:#a78bfa,color:#e2e8f0
Component Vai trò Min HA
MON (Monitor) Cluster map, quorum, authentication 3 (odd number)
MGR (Manager) Metrics, dashboard, orchestration 2 (active-standby)
OSD (Object Storage Daemon) Lưu trữ data thực tế, replication 3+ (1 per disk)
MDS (Metadata Server) Metadata cho CephFS (chỉ cần khi dùng CephFS) 2 (active-standby)
RGW (RADOS Gateway) S3/Swift API cho object storage 2+ (behind LB)

1.2. CRUSH Algorithm

graph TD
    subgraph STEP1["1️⃣ Object → Pool → PG"]
        OBJ["📄 photo.jpg"] -->|"hash() mod num_PGs"| PG["PG 3.1a"]
    end

    subgraph STEP2["2️⃣ CRUSH Map quyết định OSD set"]
        PG2["PG 3.1a"] -->|"CRUSH()"| OSDSET["OSD Set"]
    end

    subgraph STEP3["3️⃣ Replicate tới 3 OSDs"]
        OSD5["💾 osd.5<br/>PRIMARY<br/>worker1"]
        OSD12["💾 osd.12<br/>REPLICA<br/>worker2"]
        OSD8["💾 osd.8<br/>REPLICA<br/>worker3"]
    end

    PG --> PG2
    OSDSET --> OSD5
    OSDSET --> OSD12
    OSDSET --> OSD8

    style STEP1 fill:#0f172a,stroke:#3b82f6,color:#e2e8f0
    style STEP2 fill:#1e3a5f,stroke:#60a5fa,color:#e2e8f0
    style STEP3 fill:#15803d,stroke:#4ade80,color:#e2e8f0
    style OSD5 fill:#15803d,stroke:#22c55e,color:#e2e8f0
    style OSD12 fill:#1e3a5f,stroke:#3b82f6,color:#e2e8f0
    style OSD8 fill:#1e3a5f,stroke:#3b82f6,color:#e2e8f0

✅ Không cần lookup table → Scale tới exabytes ✅ Failure domain aware (rack, host, datacenter) ✅ Client tính toán trực tiếp vị trí data


PHẦN 2: ROOK — CEPH OPERATOR CHO KUBERNETES

2.1. Rook Architecture

graph TB
    subgraph K8S["☸ Kubernetes Cluster"]
        subgraph OPERATOR["🔧 Rook Operator"]
            op["Deployment<br/>• Watches CephCluster CRD<br/>• Manages Ceph daemons as K8s pods<br/>• Auto-healing, scaling"]
        end

        subgraph DAEMONS["Ceph Daemons as Pods"]
            MON["🔍 ceph-mon<br/>3 replicas"]
            MGR["📊 ceph-mgr<br/>2 replicas<br/>+ Dashboard"]
            OSD["💾 ceph-osd<br/>1 per disk"]
        end

        subgraph CRDS["📋 Custom Resource Definitions"]
            CR1["CephCluster"]
            CR2["CephBlockPool"]
            CR3["CephFilesystem"]
            CR4["CephObjectStore"]
        end
    end

    op -->|manages| MON
    op -->|manages| MGR
    op -->|manages| OSD
    op -->|watches| CR1
    CR2 -.-> OSD
    CR3 -.-> OSD
    CR4 -.-> OSD

    style K8S fill:#0f172a,stroke:#3b82f6,color:#e2e8f0
    style OPERATOR fill:#7c3aed,stroke:#a78bfa,color:#e2e8f0
    style DAEMONS fill:#1e3a5f,stroke:#60a5fa,color:#e2e8f0
    style CRDS fill:#1e293b,stroke:#475569,color:#e2e8f0

2.2. So sánh Storage Solutions

Tiêu chí Rook-Ceph Longhorn OpenEBS local-path
Block storage ✅ RBD ✅ ✅ ❌
Filesystem (RWX) ✅ CephFS ❌ (NFS hack) ❌ ❌
Object (S3) ✅ RGW ❌ ❌ ❌
Replication 3-way 3-way 3-way ❌
Performance Xuất sắc Tốt Tốt Best (local)
Complexity Cao Thấp Trung bình Rất thấp
Min nodes 3 3 3 1
CNCF Graduated Incubating Sandbox -
Best for Production Small clusters Dev/test Single node

👉 Chọn Rook-Ceph cho production on-premises: unified storage (Block + FS + Object), proven at scale, CNCF Graduated.


PHẦN 3: PLANNING HARDWARE CHO CEPH

3.1. OSD Node Sizing

graph LR
    subgraph OSD_REQ["💾 Per-OSD Requirements"]
        CPU["🔧 CPU: 1-2 cores"]
        RAM["🧠 RAM: 5GB bluestore"]
        NVME["⚡ NVMe: WAL + DB"]
        NET["🌐 Network: 10Gbps"]
    end

    subgraph EXAMPLE["📋 worker1: 4 data SSDs"]
        E1["4 OSDs × 5GB = 20GB RAM"]
        E2["4 OSDs × 2 cores = 8 CPU"]
        E3["+ workload riêng"]
    end

    OSD_REQ --> EXAMPLE

    style OSD_REQ fill:#1e3a5f,stroke:#3b82f6,color:#e2e8f0
    style EXAMPLE fill:#0f172a,stroke:#f59e0b,color:#e2e8f0

3.2. Capacity Planning

# Ceph usable capacity formula:
# Usable = Raw Capacity × (1 / Replication Factor) × Pool Utilization Target

Ví dụ:

3 workers × 4 SSDs × 500GB = 6,000 GB Raw

Replication Factor = 3 (data replicate 3 bản)

Usable = 6,000 / 3 = 2,000 GB

⚠️ Ceph khuyến nghị KHÔNG vượt 85% → 2,000 × 0.85 = 1,700 GB usable

Cho lab (mỗi worker 1 SSD 100GB):

3 × 100GB = 300GB Raw → 100GB usable → 85GB safe

3.3. Network Planning

graph LR
    subgraph NODE1["💻 OSD Node 1"]
        P1["🌐 Public Net<br/>10.10.20.x<br/>client I/O"]
        C1["🔄 Cluster Net<br/>10.10.30.x<br/>replication"]
    end

    subgraph NODE2["💻 OSD Node 2"]
        P2["🌐 Public Net<br/>10.10.20.x"]
        C2["🔄 Cluster Net<br/>10.10.30.x"]
    end

    subgraph NODE3["💻 OSD Node 3"]
        P3["🌐 Public Net<br/>10.10.20.x"]
        C3["🔄 Cluster Net<br/>10.10.30.x"]
    end

    P1 <-->|"Client I/O"| P2
    P2 <-->|"Client I/O"| P3
    C1 <-->|"Replication"| C2
    C2 <-->|"Replication"| C3

    style NODE1 fill:#0f172a,stroke:#3b82f6,color:#e2e8f0
    style NODE2 fill:#0f172a,stroke:#3b82f6,color:#e2e8f0
    style NODE3 fill:#0f172a,stroke:#3b82f6,color:#e2e8f0
    style P1 fill:#1e3a5f,stroke:#60a5fa,color:#e2e8f0
    style P2 fill:#1e3a5f,stroke:#60a5fa,color:#e2e8f0
    style P3 fill:#1e3a5f,stroke:#60a5fa,color:#e2e8f0
    style C1 fill:#7c3aed,stroke:#a78bfa,color:#e2e8f0
    style C2 fill:#7c3aed,stroke:#a78bfa,color:#e2e8f0
    style C3 fill:#7c3aed,stroke:#a78bfa,color:#e2e8f0

⚠️ Cluster network: 10Gbps minimum để recovery không impact client I/O


PHẦN 4: 3 LOẠI STORAGE

4.1. Block Storage (RBD)

graph LR
    POD["🟢 Pod"] -->|mount| PVC["📋 PVC<br/>ReadWriteOnce"]
    PVC -->|provision| RBD["💿 Ceph RBD<br/>thin provisioned"]

    style POD fill:#15803d,stroke:#22c55e,color:#e2e8f0
    style PVC fill:#1e3a5f,stroke:#3b82f6,color:#e2e8f0
    style RBD fill:#7c3aed,stroke:#a78bfa,color:#e2e8f0

✅ Database (PostgreSQL, MySQL) · ✅ Stateful applications · ✅ High IOPS ⚠️ ReadWriteOnce — chỉ 1 pod mount cùng lúc

4.2. Filesystem Storage (CephFS)

graph LR
    P1["🟢 Pod 1"] --> PVC["📋 PVC<br/>ReadWriteMany"]
    P2["🟢 Pod 2"] --> PVC
    P3["🟢 Pod 3"] --> PVC
    PVC -->|mount| CEPHFS["📁 CephFS<br/>POSIX filesystem"]

    style P1 fill:#15803d,stroke:#22c55e,color:#e2e8f0
    style P2 fill:#15803d,stroke:#22c55e,color:#e2e8f0
    style P3 fill:#15803d,stroke:#22c55e,color:#e2e8f0
    style PVC fill:#1e3a5f,stroke:#3b82f6,color:#e2e8f0
    style CEPHFS fill:#7c3aed,stroke:#a78bfa,color:#e2e8f0

✅ Shared file storage (nhiều pods cùng đọc/ghi) · ✅ Content management · ✅ AI/ML training data

4.3. Object Storage (RGW)

graph LR
    APP["🟢 App"] -->|"PUT/GET/DELETE"| S3["🌐 S3 API<br/>s3://bucket/key"]
    S3 --> RGW["☁️ Ceph Object Store<br/>RADOS Gateway"]

    style APP fill:#15803d,stroke:#22c55e,color:#e2e8f0
    style S3 fill:#1e3a5f,stroke:#3b82f6,color:#e2e8f0
    style RGW fill:#7c3aed,stroke:#a78bfa,color:#e2e8f0

✅ Backup storage · ✅ Log archives · ✅ Thay thế MinIO/AWS S3


💡 KEY TAKEAWAYS

  1. Ceph = unified storage: Block + Filesystem + Object trong 1 cluster
  2. CRUSH algorithm: tính toán vị trí data không cần lookup table → scale vô hạn
  3. Rook Operator: quản lý Ceph lifecycle trên K8s bằng CRDs
  4. 5GB RAM per OSD: plan memory cẩn thận cho storage nodes
  5. 2 networks: Public (client) + Cluster (replication) để tách traffic
  6. Replication 3x: Raw capacity / 3 = usable capacity

🎯 BÀI TẬP

Bài tập 1: Capacity Planning

  • Tính usable capacity cho: 5 nodes × 6 SSDs × 1TB, replication=3
  • Tính RAM cần thiết cho Ceph (mỗi OSD 5GB)
  • Vẽ network topology cho Ceph cluster

Bài tập 2: Chuẩn bị disks

  • Xác định disk chưa sử dụng trên worker nodes: lsblk
  • Verify disk không có partition: wipefs -a /dev/sdX
  • Chuẩn bị sẵn disk cho Bài 12

📚 BÀI TIẾP THEO

Trong Bài 12: Cài đặt Rook-Ceph Operator và CephCluster, chúng ta sẽ deploy Rook Operator và tạo CephCluster trên K8s.