🎯 LESSON OBJECTIVE__HTMLTAG_68___
After completing this lesson, you will:
- ✅ Understand Ceph architecture: RADOS, OSD, MON, MDS, MGR
- ✅ Understand CRUSH algorithm and data placement
- ✅ Compare Rook-Ceph vs Longhorn vs OpenEBS vs local-path
- ✅ Planning hardware and network for Ceph cluster
- ✅ Understand 3 types of storage: Block (RBD), Filesystem (CephFS), Object (RGW)
PART 1: CEPH ARCHITECTURE OVERVIEW
1.1. Ceph
componentsgraph TB
subgraph CLIENT["🖥️ CLIENT ACCESS"]
RBD["💿 RBD<br/>Block Storage"]
CephFS["📁 CephFS<br/>Filesystem"]
RGW["🌐 RADOS Gateway<br/>Object / S3-compat"]
end
subgraph LIB["📚 LIBRADOS API"]
librados["Unified Storage API"]
end
subgraph RADOS["⚙️ RADOS LAYER"]
MON["🔍 MON<br/>Monitor<br/>3 nodes"]
MGR["📊 MGR<br/>Manager<br/>2 nodes"]
MDS["📂 MDS<br/>Metadata Server<br/>CephFS only"]
OSD["💾 OSD<br/>Object Storage<br/>Daemon"]
end
subgraph CRUSH["🗺️ CRUSH MAP"]
crush_algo["Data placement algorithm — no lookup table!"]
end
RBD --> librados
CephFS --> librados
RGW --> librados
librados --> MON
librados --> MGR
librados --> MDS
librados --> OSD
OSD --> crush_algo
style CLIENT fill:#0f172a,stroke:#3b82f6,color:#e2e8f0
style LIB fill:#1e3a5f,stroke:#60a5fa,color:#e2e8f0
style RADOS fill:#1e293b,stroke:#3b82f6,color:#e2e8f0
style CRUSH fill:#7c3aed,stroke:#a78bfa,color:#e2e8f0
```<!--kg-card-begin: html-->
<table>
<thead>
<tr>
<th>Component</th>
<th>Role</th>
<th>Min HA</th>
</tr>
</thead>
<tbody>
<tr>
<td>MON (Monitor)</td>
<td>Cluster map, quorum, authentication</td>
<td>3 (odd number)</td>
</tr>
<tr>
<td>MGR (Manager)</td>
<td>Metrics, dashboard, orchestration</td>
<td>2 (active-standby)</td>
</tr>
<tr>
<td>OSD (Object Storage Daemon)</td>
<td>Actual data storage, replication</td>
<td>3+ (1 per disk)</td>
</tr>
<tr>
<td>MDS (Metadata Server)</td>
<td>Metadata for CephFS (only needed when using CephFS)</td>
<td>2 (active-standby)</td>
</tr>
<tr>
<td>RGW (RADOS Gateway)</td>
<td>S3/Swift API for object storage</td>
<td>2+ (behind LB)</td>
</tr>
</tbody>
</table>
<!--kg-card-end: html-->
<h3 id="12-crush-algorithm">1.2. CRUSH Algorithm</h3>
```mermaid
graph TD
subgraph STEP1["1️⃣ Object → Pool → PG"]
OBJ["📄 photo.jpg"] -->|"hash() mod num_PGs"| PG["PG 3.1a"]
end
subgraph STEP2["2️⃣ CRUSH Map quyết định OSD set"]
PG2["PG 3.1a"] -->|"CRUSH()"| OSDSET["OSD Set"]
end
subgraph STEP3["3️⃣ Replicate tới 3 OSDs"]
OSD5["💾 osd.5<br/>PRIMARY<br/>worker1"]
OSD12["💾 osd.12<br/>REPLICA<br/>worker2"]
OSD8["💾 osd.8<br/>REPLICA<br/>worker3"]
end
PG --> PG2
OSDSET --> OSD5
OSDSET --> OSD12
OSDSET --> OSD8
style STEP1 fill:#0f172a,stroke:#3b82f6,color:#e2e8f0
style STEP2 fill:#1e3a5f,stroke:#60a5fa,color:#e2e8f0
style STEP3 fill:#15803d,stroke:#4ade80,color:#e2e8f0
style OSD5 fill:#15803d,stroke:#22c55e,color:#e2e8f0
style OSD12 fill:#1e3a5f,stroke:#3b82f6,color:#e2e8f0
style OSD8 fill:#1e3a5f,stroke:#3b82f6,color:#e2e8f0
✅ No need to lookup table → Scale to exabytes ✅ Failure domain aware (rack, host, datacenter) ✅ Client directly calculates data location
PART 2: ROOK — CEPH OPERATOR FOR KUBERNETES
2.1. Rook Architecture
graph TB
subgraph K8S["☸ Kubernetes Cluster"]
subgraph OPERATOR["🔧 Rook Operator"]
op["Deployment<br/>• Watches CephCluster CRD<br/>• Manages Ceph daemons as K8s pods<br/>• Auto-healing, scaling"]
end
subgraph DAEMONS["Ceph Daemons as Pods"]
MON["🔍 ceph-mon<br/>3 replicas"]
MGR["📊 ceph-mgr<br/>2 replicas<br/>+ Dashboard"]
OSD["💾 ceph-osd<br/>1 per disk"]
end
subgraph CRDS["📋 Custom Resource Definitions"]
CR1["CephCluster"]
CR2["CephBlockPool"]
CR3["CephFilesystem"]
CR4["CephObjectStore"]
end
end
op -->|manages| MON
op -->|manages| MGR
op -->|manages| OSD
op -->|watches| CR1
CR2 -.-> OSD
CR3 -.-> OSD
CR4 -.-> OSD
style K8S fill:#0f172a,stroke:#3b82f6,color:#e2e8f0
style OPERATOR fill:#7c3aed,stroke:#a78bfa,color:#e2e8f0
style DAEMONS fill:#1e3a5f,stroke:#60a5fa,color:#e2e8f0
style CRDS fill:#1e293b,stroke:#475569,color:#e2e8f0
2.2. Compare Storage Solutions
| Criteria | Rook-Ceph | Longhorn | OpenEBS | local-path |
|---|---|---|---|---|
| Block storage | ✅ RBD | ✅ | ✅ | ❌ |
| Filesystem (RWX) | ✅ CephFS | ❌ (NFS hack) | ❌ | ❌ |
| Object (S3) | ✅ RGW | ❌ | ❌ | ❌ |
| Replication | 3-way | 3-way | 3-way | ❌ |
| Performance | Excellent__HTMLTAG_222___ | Good | Good | Best (local) |
| Complexity | High | Low | Average | Very low |
| Min nodes | 3 | 3 | 3 | 1 |
| CNCF | Graduated | Incubating | Sandbox | - |
| Best for | Production | Small clusters | Dev/test | Single node |
👉 Select Rook-Ceph for production on-premises: unified storage (Block + FS + Object), proven at scale, CNCF Graduated.
PART 3: PLANNING HARDWARE FOR CEPH
3.1. OSD Node Sizing
graph LR
subgraph OSD_REQ["💾 Per-OSD Requirements"]
CPU["🔧 CPU: 1-2 cores"]
RAM["🧠 RAM: 5GB bluestore"]
NVME["⚡ NVMe: WAL + DB"]
NET["🌐 Network: 10Gbps"]
end
subgraph EXAMPLE["📋 worker1: 4 data SSDs"]
E1["4 OSDs × 5GB = 20GB RAM"]
E2["4 OSDs × 2 cores = 8 CPU"]
E3["+ workload riêng"]
end
OSD_REQ --> EXAMPLE
style OSD_REQ fill:#1e3a5f,stroke:#3b82f6,color:#e2e8f0
style EXAMPLE fill:#0f172a,stroke:#f59e0b,color:#e2e8f0
3.2. Capacity Planning
# Ceph usable capacity formula: # Usable = Raw Capacity × (1 / Replication Factor) × Pool Utilization TargetVí dụ:
3 workers × 4 SSDs × 500GB = 6,000 GB Raw
Replication Factor = 3 (data replicate 3 bản)
Usable = 6,000 / 3 = 2,000 GB
⚠️ Ceph khuyến nghị KHÔNG vượt 85% → 2,000 × 0.85 = 1,700 GB usable
Cho lab (mỗi worker 1 SSD 100GB):
3 × 100GB = 300GB Raw → 100GB usable → 85GB safe
3.3. Network Planning
graph LR
subgraph NODE1["💻 OSD Node 1"]
P1["🌐 Public Net<br/>10.10.20.x<br/>client I/O"]
C1["🔄 Cluster Net<br/>10.10.30.x<br/>replication"]
end
subgraph NODE2["💻 OSD Node 2"]
P2["🌐 Public Net<br/>10.10.20.x"]
C2["🔄 Cluster Net<br/>10.10.30.x"]
end
subgraph NODE3["💻 OSD Node 3"]
P3["🌐 Public Net<br/>10.10.20.x"]
C3["🔄 Cluster Net<br/>10.10.30.x"]
end
P1 <-->|"Client I/O"| P2
P2 <-->|"Client I/O"| P3
C1 <-->|"Replication"| C2
C2 <-->|"Replication"| C3
style NODE1 fill:#0f172a,stroke:#3b82f6,color:#e2e8f0
style NODE2 fill:#0f172a,stroke:#3b82f6,color:#e2e8f0
style NODE3 fill:#0f172a,stroke:#3b82f6,color:#e2e8f0
style P1 fill:#1e3a5f,stroke:#60a5fa,color:#e2e8f0
style P2 fill:#1e3a5f,stroke:#60a5fa,color:#e2e8f0
style P3 fill:#1e3a5f,stroke:#60a5fa,color:#e2e8f0
style C1 fill:#7c3aed,stroke:#a78bfa,color:#e2e8f0
style C2 fill:#7c3aed,stroke:#a78bfa,color:#e2e8f0
style C3 fill:#7c3aed,stroke:#a78bfa,color:#e2e8f0
⚠️ Cluster network: 10Gbps minimum so recovery does not impact client I/O
PART 4: 3 TYPES OF STORAGE
4.1. Block Storage (RBD)
graph LR
POD["🟢 Pod"] -->|mount| PVC["📋 PVC<br/>ReadWriteOnce"]
PVC -->|provision| RBD["💿 Ceph RBD<br/>thin provisioned"]
style POD fill:#15803d,stroke:#22c55e,color:#e2e8f0
style PVC fill:#1e3a5f,stroke:#3b82f6,color:#e2e8f0
style RBD fill:#7c3aed,stroke:#a78bfa,color:#e2e8f0
✅ Database (PostgreSQL, MySQL) · ✅ Stateful applications · ✅ High IOPS ⚠️ ReadWriteOnce — only 1 pod mount at a time
4.2. Filesystem Storage (CephFS)
graph LR
P1["🟢 Pod 1"] --> PVC["📋 PVC<br/>ReadWriteMany"]
P2["🟢 Pod 2"] --> PVC
P3["🟢 Pod 3"] --> PVC
PVC -->|mount| CEPHFS["📁 CephFS<br/>POSIX filesystem"]
style P1 fill:#15803d,stroke:#22c55e,color:#e2e8f0
style P2 fill:#15803d,stroke:#22c55e,color:#e2e8f0
style P3 fill:#15803d,stroke:#22c55e,color:#e2e8f0
style PVC fill:#1e3a5f,stroke:#3b82f6,color:#e2e8f0
style CEPHFS fill:#7c3aed,stroke:#a78bfa,color:#e2e8f0
✅ Shared file storage (multiple pods read/write together) · ✅ Content management · ✅ AI/ML training data
4.3. Object Storage (RGW)
graph LR
APP["🟢 App"] -->|"PUT/GET/DELETE"| S3["🌐 S3 API<br/>s3://bucket/key"]
S3 --> RGW["☁️ Ceph Object Store<br/>RADOS Gateway"]
style APP fill:#15803d,stroke:#22c55e,color:#e2e8f0
style S3 fill:#1e3a5f,stroke:#3b82f6,color:#e2e8f0
style RGW fill:#7c3aed,stroke:#a78bfa,color:#e2e8f0
✅ Backup storage · ✅ Log archives · ✅ MinIO/AWS S3 replacement
💡 KEY TAKEAWAYS
- Ceph = unified storage: Block + Filesystem + Object in 1 cluster
- CRUSH algorithm: calculate data position without lookup table → scale infinite__HTMLTAG_314___
- Rook Operator: manage Ceph lifecycle on K8s with CRDs
- 5GB RAM per OSD: plan memory carefully for storage nodes
- 2 networks: Public (client) + Cluster (replication) to separate traffic
- Replication 3x: Raw capacity / 3 = usable capacity
🎯 EXERCISE
Exercise 1: Capacity Planning
- Calculate usable capacity for: 5 nodes × 6 SSDs × 1TB, replication=3
- Calculate RAM required for Ceph (5GB OSD each)
- Draw network topology for Ceph cluster
Exercise 2: Prepare disks__HTMLTAG_346___
- Identify unused disk on worker nodes:
lsblk
- Verify disk without partition:
wipefs -a /dev/sdX
- Prepare the disk for Lesson 12
lsblkwipefs -a /dev/sdX📚 NEXT POST
In Lesson 12: Installing Rook-Ceph Operator and CephCluster, we will deploy Rook Operator and create CephCluster on K8s.