🎯 MỤC TIÊU BÀI HỌC
Sau khi hoàn thành bài học này, bạn sẽ:
- ✅ Tính toán được sizing chính xác cho từng loại node (Control Plane, Worker, Storage)
- ✅ Thiết kế được network topology production-grade với VLAN separation
- ✅ Cấu hình được NIC bonding cho HA networking
- ✅ Hiểu và chọn được MTU phù hợp cho từng network segment
- ✅ Lập được bảng kế hoạch phần cứng hoàn chỉnh cho dự án thực tế
PHẦN 1: SIZING CONTROL PLANE NODES
1.1. Các thành phần chạy trên Control Plane
Control Plane Node
├── kube-apiserver ─── API endpoint, xử lý tất cả requests
├── etcd ─── Distributed KV store (cluster state)
├── kube-scheduler ─── Pod scheduling decisions
├── kube-controller-manager ─── Reconciliation loops
├── cloud-controller-manager ─── (Không dùng cho on-prem)
├── kubelet ─── Node agent
├── containerd ─── Container runtime
└── Cilium agent ─── CNI networking
1.2. Tính toán Resources cho Control Plane
etcd là thành phần critical nhất
etcd performance phụ thuộc chủ yếu vào disk I/O. Đây là sizing guidelines từ etcd documentation:
| Cluster Size | Nodes | Pods | etcd CPU | etcd RAM | etcd Disk | Disk Type |
|---|---|---|---|---|---|---|
| Small | < 10 | < 500 | 2 cores | 4GB | 50GB | SSD |
| Medium | 10-50 | 500-5000 | 4 cores | 8GB | 100GB | NVMe SSD |
| Large | 50-100 | 5000+ | 8 cores | 16GB | 200GB | NVMe SSD |
⚠️ Critical: etcd yêu cầu disk latency p99 < 10ms. Dùng NVMe SSD chuyên dụng cho etcd.
kube-apiserver sizing
# Tính toán dựa trên số lượng requests/giây API Server resources = f(number_of_nodes, number_of_pods, number_of_controllers)Baseline (10 nodes, 500 pods): CPU: 2 cores RAM: 4GB
Scaling rule: +1 CPU per 1000 pods +2GB RAM per 1000 pods +1 CPU per 20 nodes
Tổng hợp Control Plane Node Sizing
| Component | CPU Request | CPU Limit | RAM Request | RAM Limit |
|---|---|---|---|---|
| kube-apiserver | 250m | 2000m | 512Mi | 4Gi |
| etcd | 500m | 4000m | 1Gi | 8Gi |
| kube-scheduler | 100m | 500m | 128Mi | 512Mi |
| kube-controller-manager | 200m | 1000m | 256Mi | 1Gi |
| kubelet + containerd | 200m | 500m | 256Mi | 1Gi |
| Cilium agent | 100m | 500m | 256Mi | 1Gi |
| OS overhead | 500m | - | 1Gi | - |
| TỔNG | ~2 cores | ~8 cores | ~3.5Gi | ~16Gi |
💡 Recommendation cho Production:
Control Plane Node (mỗi node):
CPU: 8 cores (headroom cho burst)
RAM: 16GB (minimum) - 32GB (recommended)
Disk: 100GB NVMe SSD (OS + etcd)
→ etcd nên trên partition/disk riêng nếu có thể
NIC: 2× 10GbE (bonding) hoặc 1× 25GbE
PHẦN 2: SIZING WORKER NODES
2.1. Tính toán từ Workload Requirements
Công thức tính Worker nodes:
Total Worker Resources = Σ (all pod requests) + System Reserved + BufferVí dụ với hệ thống 20 microservices: ┌──────────────────────────────────────────────────────────────┐ │ Microservices (20 services × 2 replicas × avg 500m/1Gi): │ │ CPU: 20 × 2 × 500m = 20,000m = 20 cores │ │ RAM: 20 × 2 × 1Gi = 40Gi │ │ │ │ Databases (PostgreSQL 3 nodes, Redis 3, RabbitMQ 3, Kafka 3):│ │ CPU: 12 × 2000m = 24,000m = 24 cores │ │ RAM: 12 × 4Gi = 48Gi │ │ │ │ Observability (Prometheus×2, Grafana, Loki, Tempo, Alloy): │ │ CPU: ~8 cores │ │ RAM: ~24Gi │ │ │ │ Platform (ArgoCD, Vault, Istio, cert-manager, Kyverno): │ │ CPU: ~6 cores │ │ RAM: ~16Gi │ │ │ │ System Reserved per node (kubelet, containerd, Cilium, OS): │ │ CPU: ~1.5 cores × N nodes │ │ RAM: ~2Gi × N nodes │ │ │ │ TỔNG REQUEST: │ │ CPU: ~58 cores + (1.5 × N) │ │ RAM: ~128Gi + (2 × N) │ └──────────────────────────────────────────────────────────────┘
Buffer (30% headroom cho HA + burst): CPU: 58 × 1.3 = ~76 cores RAM: 128 × 1.3 = ~167Gi
Sizing calculation: Nếu mỗi worker: 16 cores, 64GB RAM Số workers = max(76/16, 167/64) = max(4.75, 2.6) = 5 workers
💡 Khuyến nghị: 5-6 worker nodes × (16 cores, 64GB RAM) → Cho phép mất 1 node mà workloads vẫn schedulable
2.2. Sizing Guidelines theo Quy mô
| Quy mô | Services | Workers | CPU/node | RAM/node | Disk/node |
|---|---|---|---|---|---|
| Small (Lab) | 5-10 | 3 | 8 cores | 32GB | 200GB SSD |
| Medium | 10-30 | 5-8 | 16 cores | 64GB | 500GB NVMe |
| Large | 30-100 | 10-20 | 32 cores | 128GB | 1TB NVMe |
| XLarge | 100+ | 20+ | 64 cores | 256GB | 2TB NVMe |
2.3. System Reserved Resources
Kubernetes cần reserve resources cho system components trên mỗi node:
# /var/lib/kubelet/config.yaml
apiVersion: kubelet.config.k8s.io/v1beta1
kind: KubeletConfiguration
systemReserved:
cpu: "500m"
memory: "1Gi"
ephemeral-storage: "10Gi"
kubeReserved:
cpu: "500m"
memory: "1Gi"
ephemeral-storage: "5Gi"
evictionHard:
memory.available: "500Mi"
nodefs.available: "10%"
imagefs.available: "15%"
Allocatable = Total - systemReserved - kubeReserved - evictionThreshold
Ví dụ: Worker node 16 cores, 64GB RAM:
CPU allocatable: 16000m - 500m - 500m = 15000m
RAM allocatable: 64Gi - 1Gi - 1Gi - 500Mi = 61.5Gi
PHẦN 3: SIZING STORAGE NODES (CEPH)
3.1. Ceph Components trên Storage Nodes
Storage Node
├── Ceph OSD daemon (1 per disk) ─── Object Storage Daemon
│ ├── BlueStore (direct disk I/O)
│ └── WAL + DB on SSD (nếu dùng HDD)
├── Ceph MON (trên 3 nodes) ─── Cluster monitor
├── Ceph MGR (trên 2 nodes) ─── Manager, dashboard
└── kubelet + containerd + Cilium ─── K8s agent
3.2. Tính toán Storage Capacity
Usable Capacity = Raw Capacity / Replication Factor × Utilization TargetVí dụ: 3 nodes × 4 disks × 2TB = 24TB raw Replication factor = 3 (data replicated 3 lần) Utilization target = 75% (để headroom cho recovery)
Usable = 24TB / 3 × 0.75 = 6TB usable
Phân bổ: PostgreSQL data: 500GB (× 3 replicas nguồn PG) Kafka log retention: 500GB Loki logs: 1TB Thanos metrics: 500GB Velero backups: 1TB Application data: 500GB Buffer: 2TB ───────────────────────────── TỔNG: ~6TB ❯ Khớp 6TB usable
3.3. Sizing Ceph OSD Nodes
| Component | Sizing Rule | Ví dụ (4 OSDs/node) |
|---|---|---|
| CPU per OSD | 1 core per OSD (min) | 4 cores cho OSDs |
| RAM per OSD | 5GB per OSD (BlueStore default) | 20GB cho OSDs |
| Ceph MON RAM | ~2-4GB | 4GB |
| System + K8s | ~4GB RAM, 2 cores | 4GB, 2 cores |
| TỔNG per node | 6 cores, 28GB RAM |
💡 Recommendation:
Storage Node (dedicated hoặc converged với worker):
CPU: 8 cores
RAM: 32-64GB (phụ thuộc số OSD)
Disk: 1× 500GB NVMe (OS)
4× 2TB NVMe (Ceph OSD)
NIC: 2× 25GbE (1 cluster + 1 public network)
⚠️ Quyết định quan trọng: Dedicated storage nodes vs Converged (worker + storage)?
Dedicated Storage Nodes: ✅ Isolation: Storage I/O không ảnh hưởng workloads ✅ Independent scaling ❌ Thêm serversConverged (Worker + Storage cùng node): ✅ Ít servers, tận dụng hardware ❌ Noisy neighbor: Ceph I/O có thể ảnh hưởng pods ❌ Node failure mất cả compute + storage
→ Production: Dedicated storage nodes → Lab/Small: Converged OK
PHẦN 4: NETWORK TOPOLOGY DESIGN
4.1. 4 Networks cho Production
┌─────────────────────────────────────────────────────────────────┐
│ NETWORK TOPOLOGY │
├─────────────────────────────────────────────────────────────────┤
│ │
│ ┌── VLAN 10: Management Network (192.168.10.0/24) ──────────┐│
│ │ SSH access, monitoring, IPMI/iDRAC/iLO ││
│ │ MTU: 1500 ││
│ │ NIC: eth0 (hoặc bond0 member) ││
│ └────────────────────────────────────────────────────────────┘│
│ │
│ ┌── VLAN 20: Cluster Network (10.10.20.0/24) ───────────────┐│
│ │ K8s API, Pod-to-Pod traffic, Service communication ││
│ │ MTU: 9000 (Jumbo Frames) ││
│ │ NIC: bond0 (eth1 + eth2 LACP) ││
│ └────────────────────────────────────────────────────────────┘│
│ │
│ ┌── VLAN 30: Storage Network (10.10.30.0/24) ───────────────┐│
│ │ Ceph cluster traffic (OSD replication, recovery) ││
│ │ MTU: 9000 (Jumbo Frames) ││
│ │ NIC: bond1 (eth3 + eth4 LACP) - Dedicated 25GbE ││
│ └────────────────────────────────────────────────────────────┘│
│ │
│ ┌── VLAN 40: External Network (10.10.40.0/24) ──────────────┐│
│ │ User traffic, Ingress, MetalLB VIPs ││
│ │ MTU: 1500 ││
│ │ NIC: bond0 (shared với Cluster, VLAN tagged) ││
│ └────────────────────────────────────────────────────────────┘│
│ │
│ Firewall/Router: Giữa External ↔ Internal networks │
│ DNS: Internal DNS cho *.k8s.local │
└─────────────────────────────────────────────────────────────────┘
4.2. IP Planning chi tiết
| Node | Management (VLAN 10) | Cluster (VLAN 20) | Storage (VLAN 30) | Role |
|---|---|---|---|---|
| lb1 | 192.168.10.9 | 10.10.20.9 | - | HAProxy/keepalived primary |
| lb2 | 192.168.10.10 | 10.10.20.10 | - | HAProxy/keepalived backup |
| VIP | - | 10.10.20.100 | - | K8s API Server VIP |
| master1 | 192.168.10.11 | 10.10.20.11 | - | Control Plane 1 |
| master2 | 192.168.10.12 | 10.10.20.12 | - | Control Plane 2 |
| master3 | 192.168.10.13 | 10.10.20.13 | - | Control Plane 3 |
| worker1 | 192.168.10.21 | 10.10.20.21 | - | Worker Node 1 |
| worker2 | 192.168.10.22 | 10.10.20.22 | - | Worker Node 2 |
| worker3 | 192.168.10.23 | 10.10.20.23 | - | Worker Node 3 |
| storage1 | 192.168.10.31 | 10.10.20.31 | 10.10.30.31 | Ceph OSD Node 1 |
| storage2 | 192.168.10.32 | 10.10.20.32 | 10.10.30.32 | Ceph OSD Node 2 |
| storage3 | 192.168.10.33 | 10.10.20.33 | 10.10.30.33 | Ceph OSD Node 3 |
| MetalLB Pool | - | - | - | 10.10.40.200-250 |
4.3. NIC Bonding Configuration
NIC bonding (LACP) cung cấp link redundancy và bandwidth aggregation:
# /etc/netplan/01-bonding.yaml (Ubuntu 24.04)
network:
version: 2
renderer: networkd
ethernets:
eth0:
dhcp4: false
eth1:
dhcp4: false
eth2:
dhcp4: false
eth3:
dhcp4: false
eth4:
dhcp4: false
bonds:
bond0:
interfaces: [eth1, eth2]
parameters:
mode: 802.3ad # LACP
lacp-rate: fast
mii-monitor-interval: 100
transmit-hash-policy: layer3+4
mtu: 9000
bond1:
interfaces: [eth3, eth4]
parameters:
mode: 802.3ad
lacp-rate: fast
mii-monitor-interval: 100
transmit-hash-policy: layer3+4
mtu: 9000
vlans:
bond0.20:
id: 20
link: bond0
mtu: 9000
addresses:
- 10.10.20.21/24
routes:
- to: 10.244.0.0/16 # Pod CIDR
via: 10.10.20.1
- to: 10.96.0.0/12 # Service CIDR
via: 10.10.20.1
bond0.40:
id: 40
link: bond0
addresses:
- 10.10.40.21/24
routes:
- to: default
via: 10.10.40.1
bond1.30:
id: 30
link: bond1
mtu: 9000
addresses:
- 10.10.30.21/24
# Apply cấu hình
sudo netplan apply
# Verify bonding
cat /proc/net/bonding/bond0
# Output:
# Bonding Mode: IEEE 802.3ad Dynamic link aggregation
# MII Status: up
# Slave Interface: eth1 → MII Status: up
# Slave Interface: eth2 → MII Status: up
# Verify VLAN
ip -d link show bond0.20
# Test MTU
ping -M do -s 8972 10.10.20.11 # 8972 + 28 = 9000 MTU
4.4. MTU Sizing
| Network | MTU | Lý do |
|---|---|---|
| Management | 1500 | Standard, tương thích mọi device |
| Cluster (K8s) | 9000 | Jumbo frames giảm CPU overhead, tăng throughput |
| Storage (Ceph) | 9000 | Critical cho Ceph OSD replication performance |
| External | 1500 | Standard cho internet-facing traffic |
| Pod Network (Cilium) | 8950 | MTU underlay (9000) - VXLAN overhead (50) |
⚠️ Jumbo Frames requirement: Tất cả switch ports trên path phải support và enable MTU 9000. Kiểm tra với switch admin trước khi deploy.
PHẦN 5: SWITCH VÀ FIREWALL REQUIREMENTS
5.1. Switch Requirements
Top-of-Rack (ToR) Switch Requirements: ├── L2/L3 capable ├── VLAN support (802.1Q) ├── LACP support (802.3ad) ├── Jumbo frames (MTU 9000) ├── Spanning Tree (RSTP/MSTP) └── Port count: 24-48 × 10/25GbE + 4-8 uplinks
Recommended Models (by budget): Budget: Arista 7010T, Dell S3048-ON Mid-range: Arista 7050SX, Cisco Nexus 93180YC Enterprise: Arista 7280R, Cisco Nexus 9336C
5.2. Firewall Rules giữa Networks
| Source | Destination | Port | Protocol | Purpose |
|---|---|---|---|---|
| Management | All nodes | 22 | TCP | SSH |
| External | Worker/LB | 80, 443 | TCP | HTTP/HTTPS Ingress |
| External | Master VIP | 6443 | TCP | K8s API (nếu cần external) |
| Cluster | Cluster | 6443 | TCP | K8s API Server |
| Cluster | Cluster | 2379-2380 | TCP | etcd peer & client |
| Cluster | Cluster | 10250 | TCP | kubelet API |
| Cluster | Cluster | 10259 | TCP | kube-scheduler |
| Cluster | Cluster | 10257 | TCP | kube-controller-manager |
| Cluster | Cluster | 30000-32767 | TCP | NodePort range |
| Cluster | Cluster | 4240, 4244 | TCP | Cilium health, Hubble |
| Cluster | Cluster | 8472 | UDP | Cilium VXLAN |
| Storage | Storage | 6789 | TCP | Ceph MON |
| Storage | Storage | 6800-7300 | TCP | Ceph OSD |
| Cluster | Storage | 6789,6800-7300 | TCP | Ceph client access |
PHẦN 6: DISK LAYOUT VÀ PARTITIONING
6.1. Control Plane Disk Layout
Disk: 1× 500GB NVMe SSD ├── /boot/efi 200MB (EFI System Partition) ├── /boot 1GB (kernel, initramfs) ├── / 50GB (OS root) ├── /var/lib/etcd 100GB (etcd data - SEPARATE partition!) ├── /var/lib/containerd 100GB (container images, layers) ├── /var/log 50GB (system logs) └── (remaining) ~200GB (buffer)
LVM recommended cho flexibility: VG: vg-system → LV: lv-root, lv-etcd, lv-containerd, lv-log
6.2. Worker Node Disk Layout
Disk: 1× 500GB NVMe SSD (OS)
├── /boot/efi 200MB
├── /boot 1GB
├── / 50GB
├── /var/lib/containerd 200GB (container images!)
├── /var/log 50GB
└── (remaining) ~200GB
Raw disks cho Ceph OSD (nếu converged mode): /dev/sdb → Ceph OSD 0 /dev/sdc → Ceph OSD 1 (KHÔNG partition, KHÔNG format — Rook sẽ manage)
6.3. Storage Node Disk Layout
Disk 1: 500GB NVMe (OS) ├── / (OS root) ├── /var/lib/containerd └── /var/logDisk 2-5: 4× 2TB NVMe (Ceph OSD) /dev/nvme1n1 → Ceph OSD 0 (RAW - không format) /dev/nvme2n1 → Ceph OSD 1 (RAW) /dev/nvme3n1 → Ceph OSD 2 (RAW) /dev/nvme4n1 → Ceph OSD 3 (RAW)
⚠️ QUAN TRỌNG: KHÔNG tạo partition hay filesystem trên Ceph disks. Rook-Ceph sẽ sử dụng trực tiếp raw block devices.
PHẦN 7: BILL OF MATERIALS (BOM)
7.1. BOM cho Production (Medium - 20 microservices)
| # | Component | Qty | Specs | Role |
|---|---|---|---|---|
| 1 | Server (Control Plane) | 3 | 8C/32GB/500GB NVMe, 2×10GbE | K8s masters |
| 2 | Server (Worker) | 5 | 16C/64GB/500GB NVMe, 2×25GbE | Workload nodes |
| 3 | Server (Storage) | 3 | 8C/64GB/500GB NVMe + 4×2TB NVMe, 2×25GbE | Ceph OSD |
| 4 | Server (LB) | 2 | 4C/8GB/100GB SSD, 2×10GbE | HAProxy/keepalived |
| 5 | ToR Switch | 2 | 48×25GbE + 8×100GbE uplink | Network |
| 6 | UPS | 2 | 3kVA online double-conversion | Power protection |
| 7 | PDU | 2 | Managed, dual feed | Power distribution |
Tổng: 13 servers + 2 switches + power infrastructure
💡 KEY TAKEAWAYS
- Control Plane cần NVMe SSD cho etcd, 8+ cores, 16-32GB RAM mỗi node
- Worker sizing tính từ tổng pod requests + 30% buffer + system reserved
- Ceph storage cần 5GB RAM per OSD, raw disks không format
- 4 networks riêng biệt bằng VLAN: Management, Cluster, Storage, External
- NIC bonding (LACP) cho link redundancy, Jumbo Frames 9000 cho cluster/storage
- Disk layout: etcd cần partition riêng, Ceph cần raw block devices
🎯 BÀI TẬP
Bài tập 1: Tính toán Sizing
Cho hệ thống gồm:
- 30 microservices, mỗi service 2 replicas, avg 1 CPU/2GB RAM
- PostgreSQL 3 nodes (4 CPU/8GB RAM each)
- Kafka 3 brokers (4 CPU/8GB RAM each)
- Full observability stack
- Data retention: 90 ngày logs, 1 năm metrics
Tính: Số worker nodes, storage nodes, tổng disk capacity cần thiết.
Bài tập 2: Network Design
Vẽ network topology diagram chi tiết cho hệ thống trên, bao gồm:
- VLAN assignment cho từng network segment
- IP planning table cho tất cả nodes
- NIC bonding topology
- Firewall rules matrix
Bài tập 3: Lab Setup
- Sử dụng Vagrantfile từ Bài 1, thêm NIC cho storage network
- Cấu hình VLAN trên VMs (nếu hypervisor support)
- Test connectivity: ping, iperf3 between all networks
- Verify MTU 9000 hoạt động trên cluster network
📚 BÀI TIẾP THEO
Trong Bài 3: Chuẩn bị Linux OS và System Tuning, chúng ta sẽ cấu hình kernel parameters, tắt swap, setup NTP, SSH hardening và chuẩn bị tất cả nodes trước khi cài Kubernetes.