Chuyển đến nội dung chính

LESSON 11: DISTRIBUTED STORAGE ARCHITECTURE WITH ROOK-CEPH

Overview of Ceph architecture (RADOS, OSD, MON, MDS, MGR), why choose Rook-Ceph for K8s, compare storage solutions, planning capacity and network for Ceph cluster.

🔒 DevSecOps — Lesson 11 LESSON 11: DISTRIBUTED STORAGE ARCHITECTURE WITH ROOK-CEPH

Deploy Microservices On-Premises with Kubernetes HA

Part 3: Distributed Storage — Rook-Ceph

xdev.asia

🎯 LESSON OBJECTIVE__HTMLTAG_68___

After completing this lesson, you will:

  • ✅ Understand Ceph architecture: RADOS, OSD, MON, MDS, MGR
  • ✅ Understand CRUSH algorithm and data placement
  • ✅ Compare Rook-Ceph vs Longhorn vs OpenEBS vs local-path
  • ✅ Planning hardware and network for Ceph cluster
  • ✅ Understand 3 types of storage: Block (RBD), Filesystem (CephFS), Object (RGW)

PART 1: CEPH ARCHITECTURE OVERVIEW

1.1. Ceph

components
graph TB
    subgraph CLIENT["🖥️ CLIENT ACCESS"]
        RBD["💿 RBD<br/>Block Storage"]
        CephFS["📁 CephFS<br/>Filesystem"]
        RGW["🌐 RADOS Gateway<br/>Object / S3-compat"]
    end

    subgraph LIB["📚 LIBRADOS API"]
        librados["Unified Storage API"]
    end

    subgraph RADOS["⚙️ RADOS LAYER"]
        MON["🔍 MON<br/>Monitor<br/>3 nodes"]
        MGR["📊 MGR<br/>Manager<br/>2 nodes"]
        MDS["📂 MDS<br/>Metadata Server<br/>CephFS only"]
        OSD["💾 OSD<br/>Object Storage<br/>Daemon"]
    end

    subgraph CRUSH["🗺️ CRUSH MAP"]
        crush_algo["Data placement algorithm — no lookup table!"]
    end

    RBD --> librados
    CephFS --> librados
    RGW --> librados
    librados --> MON
    librados --> MGR
    librados --> MDS
    librados --> OSD
    OSD --> crush_algo

    style CLIENT fill:#0f172a,stroke:#3b82f6,color:#e2e8f0
    style LIB fill:#1e3a5f,stroke:#60a5fa,color:#e2e8f0
    style RADOS fill:#1e293b,stroke:#3b82f6,color:#e2e8f0
    style CRUSH fill:#7c3aed,stroke:#a78bfa,color:#e2e8f0
```<!--kg-card-begin: html-->
<table>
<thead>
<tr>
<th>Component</th>
<th>Role</th>
<th>Min HA</th>
</tr>
</thead>
<tbody>
<tr>
<td>MON (Monitor)</td>
<td>Cluster map, quorum, authentication</td>
<td>3 (odd number)</td>
</tr>
<tr>
<td>MGR (Manager)</td>
<td>Metrics, dashboard, orchestration</td>
<td>2 (active-standby)</td>
</tr>
<tr>
<td>OSD (Object Storage Daemon)</td>
<td>Actual data storage, replication</td>
<td>3+ (1 per disk)</td>
</tr>
<tr>
<td>MDS (Metadata Server)</td>
<td>Metadata for CephFS (only needed when using CephFS)</td>
<td>2 (active-standby)</td>
</tr>
<tr>
<td>RGW (RADOS Gateway)</td>
<td>S3/Swift API for object storage</td>
<td>2+ (behind LB)</td>
</tr>
</tbody>
</table>
<!--kg-card-end: html-->

<h3 id="12-crush-algorithm">1.2. CRUSH Algorithm</h3>

```mermaid
graph TD
    subgraph STEP1["1️⃣ Object → Pool → PG"]
        OBJ["📄 photo.jpg"] -->|"hash() mod num_PGs"| PG["PG 3.1a"]
    end

    subgraph STEP2["2️⃣ CRUSH Map quyết định OSD set"]
        PG2["PG 3.1a"] -->|"CRUSH()"| OSDSET["OSD Set"]
    end

    subgraph STEP3["3️⃣ Replicate tới 3 OSDs"]
        OSD5["💾 osd.5<br/>PRIMARY<br/>worker1"]
        OSD12["💾 osd.12<br/>REPLICA<br/>worker2"]
        OSD8["💾 osd.8<br/>REPLICA<br/>worker3"]
    end

    PG --> PG2
    OSDSET --> OSD5
    OSDSET --> OSD12
    OSDSET --> OSD8

    style STEP1 fill:#0f172a,stroke:#3b82f6,color:#e2e8f0
    style STEP2 fill:#1e3a5f,stroke:#60a5fa,color:#e2e8f0
    style STEP3 fill:#15803d,stroke:#4ade80,color:#e2e8f0
    style OSD5 fill:#15803d,stroke:#22c55e,color:#e2e8f0
    style OSD12 fill:#1e3a5f,stroke:#3b82f6,color:#e2e8f0
    style OSD8 fill:#1e3a5f,stroke:#3b82f6,color:#e2e8f0

✅ No need to lookup table → Scale to exabytes ✅ Failure domain aware (rack, host, datacenter) ✅ Client directly calculates data location


PART 2: ROOK — CEPH OPERATOR FOR KUBERNETES

2.1. Rook Architecture

graph TB
    subgraph K8S["☸ Kubernetes Cluster"]
        subgraph OPERATOR["🔧 Rook Operator"]
            op["Deployment<br/>• Watches CephCluster CRD<br/>• Manages Ceph daemons as K8s pods<br/>• Auto-healing, scaling"]
        end

        subgraph DAEMONS["Ceph Daemons as Pods"]
            MON["🔍 ceph-mon<br/>3 replicas"]
            MGR["📊 ceph-mgr<br/>2 replicas<br/>+ Dashboard"]
            OSD["💾 ceph-osd<br/>1 per disk"]
        end

        subgraph CRDS["📋 Custom Resource Definitions"]
            CR1["CephCluster"]
            CR2["CephBlockPool"]
            CR3["CephFilesystem"]
            CR4["CephObjectStore"]
        end
    end

    op -->|manages| MON
    op -->|manages| MGR
    op -->|manages| OSD
    op -->|watches| CR1
    CR2 -.-> OSD
    CR3 -.-> OSD
    CR4 -.-> OSD

    style K8S fill:#0f172a,stroke:#3b82f6,color:#e2e8f0
    style OPERATOR fill:#7c3aed,stroke:#a78bfa,color:#e2e8f0
    style DAEMONS fill:#1e3a5f,stroke:#60a5fa,color:#e2e8f0
    style CRDS fill:#1e293b,stroke:#475569,color:#e2e8f0

2.2. Compare Storage Solutions

Criteria Rook-Ceph Longhorn OpenEBS local-path
Block storage ✅ RBD ✅ ✅ ❌
Filesystem (RWX) ✅ CephFS ❌ (NFS hack) ❌ ❌
Object (S3) ✅ RGW ❌ ❌ ❌
Replication 3-way 3-way 3-way ❌
Performance Excellent__HTMLTAG_222___ Good Good Best (local)
Complexity High Low Average Very low
Min nodes 3 3 3 1
CNCF Graduated Incubating Sandbox -
Best for Production Small clusters Dev/test Single node

👉 Select Rook-Ceph for production on-premises: unified storage (Block + FS + Object), proven at scale, CNCF Graduated.


PART 3: PLANNING HARDWARE FOR CEPH

3.1. OSD Node Sizing

graph LR
    subgraph OSD_REQ["💾 Per-OSD Requirements"]
        CPU["🔧 CPU: 1-2 cores"]
        RAM["🧠 RAM: 5GB bluestore"]
        NVME["⚡ NVMe: WAL + DB"]
        NET["🌐 Network: 10Gbps"]
    end

    subgraph EXAMPLE["📋 worker1: 4 data SSDs"]
        E1["4 OSDs × 5GB = 20GB RAM"]
        E2["4 OSDs × 2 cores = 8 CPU"]
        E3["+ workload riêng"]
    end

    OSD_REQ --> EXAMPLE

    style OSD_REQ fill:#1e3a5f,stroke:#3b82f6,color:#e2e8f0
    style EXAMPLE fill:#0f172a,stroke:#f59e0b,color:#e2e8f0

3.2. Capacity Planning

# Ceph usable capacity formula:
# Usable = Raw Capacity × (1 / Replication Factor) × Pool Utilization Target

Ví dụ:

3 workers × 4 SSDs × 500GB = 6,000 GB Raw

Replication Factor = 3 (data replicate 3 bản)

Usable = 6,000 / 3 = 2,000 GB

⚠️ Ceph khuyến nghị KHÔNG vượt 85% → 2,000 × 0.85 = 1,700 GB usable

Cho lab (mỗi worker 1 SSD 100GB):

3 × 100GB = 300GB Raw → 100GB usable → 85GB safe

3.3. Network Planning

graph LR
    subgraph NODE1["💻 OSD Node 1"]
        P1["🌐 Public Net<br/>10.10.20.x<br/>client I/O"]
        C1["🔄 Cluster Net<br/>10.10.30.x<br/>replication"]
    end

    subgraph NODE2["💻 OSD Node 2"]
        P2["🌐 Public Net<br/>10.10.20.x"]
        C2["🔄 Cluster Net<br/>10.10.30.x"]
    end

    subgraph NODE3["💻 OSD Node 3"]
        P3["🌐 Public Net<br/>10.10.20.x"]
        C3["🔄 Cluster Net<br/>10.10.30.x"]
    end

    P1 <-->|"Client I/O"| P2
    P2 <-->|"Client I/O"| P3
    C1 <-->|"Replication"| C2
    C2 <-->|"Replication"| C3

    style NODE1 fill:#0f172a,stroke:#3b82f6,color:#e2e8f0
    style NODE2 fill:#0f172a,stroke:#3b82f6,color:#e2e8f0
    style NODE3 fill:#0f172a,stroke:#3b82f6,color:#e2e8f0
    style P1 fill:#1e3a5f,stroke:#60a5fa,color:#e2e8f0
    style P2 fill:#1e3a5f,stroke:#60a5fa,color:#e2e8f0
    style P3 fill:#1e3a5f,stroke:#60a5fa,color:#e2e8f0
    style C1 fill:#7c3aed,stroke:#a78bfa,color:#e2e8f0
    style C2 fill:#7c3aed,stroke:#a78bfa,color:#e2e8f0
    style C3 fill:#7c3aed,stroke:#a78bfa,color:#e2e8f0

⚠️ Cluster network: 10Gbps minimum so recovery does not impact client I/O


PART 4: 3 TYPES OF STORAGE

4.1. Block Storage (RBD)

graph LR
    POD["🟢 Pod"] -->|mount| PVC["📋 PVC<br/>ReadWriteOnce"]
    PVC -->|provision| RBD["💿 Ceph RBD<br/>thin provisioned"]

    style POD fill:#15803d,stroke:#22c55e,color:#e2e8f0
    style PVC fill:#1e3a5f,stroke:#3b82f6,color:#e2e8f0
    style RBD fill:#7c3aed,stroke:#a78bfa,color:#e2e8f0

✅ Database (PostgreSQL, MySQL) · ✅ Stateful applications · ✅ High IOPS ⚠️ ReadWriteOnce — only 1 pod mount at a time

4.2. Filesystem Storage (CephFS)

graph LR
    P1["🟢 Pod 1"] --> PVC["📋 PVC<br/>ReadWriteMany"]
    P2["🟢 Pod 2"] --> PVC
    P3["🟢 Pod 3"] --> PVC
    PVC -->|mount| CEPHFS["📁 CephFS<br/>POSIX filesystem"]

    style P1 fill:#15803d,stroke:#22c55e,color:#e2e8f0
    style P2 fill:#15803d,stroke:#22c55e,color:#e2e8f0
    style P3 fill:#15803d,stroke:#22c55e,color:#e2e8f0
    style PVC fill:#1e3a5f,stroke:#3b82f6,color:#e2e8f0
    style CEPHFS fill:#7c3aed,stroke:#a78bfa,color:#e2e8f0

✅ Shared file storage (multiple pods read/write together) · ✅ Content management · ✅ AI/ML training data

4.3. Object Storage (RGW)

graph LR
    APP["🟢 App"] -->|"PUT/GET/DELETE"| S3["🌐 S3 API<br/>s3://bucket/key"]
    S3 --> RGW["☁️ Ceph Object Store<br/>RADOS Gateway"]

    style APP fill:#15803d,stroke:#22c55e,color:#e2e8f0
    style S3 fill:#1e3a5f,stroke:#3b82f6,color:#e2e8f0
    style RGW fill:#7c3aed,stroke:#a78bfa,color:#e2e8f0

✅ Backup storage · ✅ Log archives · ✅ MinIO/AWS S3 replacement


💡 KEY TAKEAWAYS

  1. Ceph = unified storage: Block + Filesystem + Object in 1 cluster
  2. CRUSH algorithm: calculate data position without lookup table → scale infinite__HTMLTAG_314___
  3. Rook Operator: manage Ceph lifecycle on K8s with CRDs
  4. 5GB RAM per OSD: plan memory carefully for storage nodes
  5. 2 networks: Public (client) + Cluster (replication) to separate traffic
  6. Replication 3x: Raw capacity / 3 = usable capacity

🎯 EXERCISE

Exercise 1: Capacity Planning

  • Calculate usable capacity for: 5 nodes × 6 SSDs × 1TB, replication=3
  • Calculate RAM required for Ceph (5GB OSD each)
  • Draw network topology for Ceph cluster

Exercise 2: Prepare disks__HTMLTAG_346___
  • Identify unused disk on worker nodes: lsblk
  • Verify disk without partition: wipefs -a /dev/sdX
  • Prepare the disk for Lesson 12

📚 NEXT POST

In Lesson 12: Installing Rook-Ceph Operator and CephCluster, we will deploy Rook Operator and create CephCluster on K8s.