Chuyển đến nội dung chính

BÀI 2: LẬP KẾ HOẠCH PHẦN CỨNG VÀ NETWORK TOPOLOGY

Tính toán CPU/RAM/Disk cho control plane, worker nodes, storage nodes. Thiết kế network topology: management network, cluster network, storage network, external network. VLAN, bonding, MTU sizing cho production.

🔒 DevSecOps — Bài 2 BÀI 2: LẬP KẾ HOẠCH PHẦN CỨNG VÀ NETWORK TOPOLOGY

Deploy Microservices On-Premises với Kubernetes HA

Phần 1: Nền tảng & Thiết kế Hạ tầng On-Premises

xdev.asia

🎯 MỤC TIÊU BÀI HỌC

Sau khi hoàn thành bài học này, bạn sẽ:

  • ✅ Tính toán được sizing chính xác cho từng loại node (Control Plane, Worker, Storage)
  • ✅ Thiết kế được network topology production-grade với VLAN separation
  • ✅ Cấu hình được NIC bonding cho HA networking
  • ✅ Hiểu và chọn được MTU phù hợp cho từng network segment
  • ✅ Lập được bảng kế hoạch phần cứng hoàn chỉnh cho dự án thực tế

PHẦN 1: SIZING CONTROL PLANE NODES

1.1. Các thành phần chạy trên Control Plane


Control Plane Node
├── kube-apiserver           ─── API endpoint, xử lý tất cả requests
├── etcd                     ─── Distributed KV store (cluster state)
├── kube-scheduler           ─── Pod scheduling decisions
├── kube-controller-manager  ─── Reconciliation loops
├── cloud-controller-manager ─── (Không dùng cho on-prem)
├── kubelet                  ─── Node agent
├── containerd               ─── Container runtime
└── Cilium agent            ─── CNI networking

1.2. Tính toán Resources cho Control Plane

etcd là thành phần critical nhất

etcd performance phụ thuộc chủ yếu vào disk I/O. Đây là sizing guidelines từ etcd documentation:

Cluster Size Nodes Pods etcd CPU etcd RAM etcd Disk Disk Type
Small < 10 < 500 2 cores 4GB 50GB SSD
Medium 10-50 500-5000 4 cores 8GB 100GB NVMe SSD
Large 50-100 5000+ 8 cores 16GB 200GB NVMe SSD

⚠️ Critical: etcd yêu cầu disk latency p99 < 10ms. Dùng NVMe SSD chuyên dụng cho etcd.

kube-apiserver sizing


# Tính toán dựa trên số lượng requests/giây
API Server resources = f(number_of_nodes, number_of_pods, number_of_controllers)

Baseline (10 nodes, 500 pods): CPU: 2 cores RAM: 4GB

Scaling rule: +1 CPU per 1000 pods +2GB RAM per 1000 pods +1 CPU per 20 nodes

Tổng hợp Control Plane Node Sizing

Component CPU Request CPU Limit RAM Request RAM Limit
kube-apiserver 250m 2000m 512Mi 4Gi
etcd 500m 4000m 1Gi 8Gi
kube-scheduler 100m 500m 128Mi 512Mi
kube-controller-manager 200m 1000m 256Mi 1Gi
kubelet + containerd 200m 500m 256Mi 1Gi
Cilium agent 100m 500m 256Mi 1Gi
OS overhead 500m - 1Gi -
TỔNG ~2 cores ~8 cores ~3.5Gi ~16Gi

💡 Recommendation cho Production:


Control Plane Node (mỗi node):
  CPU:  8 cores (headroom cho burst)
  RAM:  16GB (minimum) - 32GB (recommended)
  Disk: 100GB NVMe SSD (OS + etcd)
        → etcd nên trên partition/disk riêng nếu có thể
  NIC:  2× 10GbE (bonding) hoặc 1× 25GbE

PHẦN 2: SIZING WORKER NODES

2.1. Tính toán từ Workload Requirements

Công thức tính Worker nodes:


Total Worker Resources = Σ (all pod requests) + System Reserved + Buffer

Ví dụ với hệ thống 20 microservices: ┌──────────────────────────────────────────────────────────────┐ │ Microservices (20 services × 2 replicas × avg 500m/1Gi): │ │ CPU: 20 × 2 × 500m = 20,000m = 20 cores │ │ RAM: 20 × 2 × 1Gi = 40Gi │ │ │ │ Databases (PostgreSQL 3 nodes, Redis 3, RabbitMQ 3, Kafka 3):│ │ CPU: 12 × 2000m = 24,000m = 24 cores │ │ RAM: 12 × 4Gi = 48Gi │ │ │ │ Observability (Prometheus×2, Grafana, Loki, Tempo, Alloy): │ │ CPU: ~8 cores │ │ RAM: ~24Gi │ │ │ │ Platform (ArgoCD, Vault, Istio, cert-manager, Kyverno): │ │ CPU: ~6 cores │ │ RAM: ~16Gi │ │ │ │ System Reserved per node (kubelet, containerd, Cilium, OS): │ │ CPU: ~1.5 cores × N nodes │ │ RAM: ~2Gi × N nodes │ │ │ │ TỔNG REQUEST: │ │ CPU: ~58 cores + (1.5 × N) │ │ RAM: ~128Gi + (2 × N) │ └──────────────────────────────────────────────────────────────┘

Buffer (30% headroom cho HA + burst): CPU: 58 × 1.3 = ~76 cores RAM: 128 × 1.3 = ~167Gi

Sizing calculation: Nếu mỗi worker: 16 cores, 64GB RAM Số workers = max(76/16, 167/64) = max(4.75, 2.6) = 5 workers

💡 Khuyến nghị: 5-6 worker nodes × (16 cores, 64GB RAM) → Cho phép mất 1 node mà workloads vẫn schedulable

2.2. Sizing Guidelines theo Quy mô

Quy mô Services Workers CPU/node RAM/node Disk/node
Small (Lab) 5-10 3 8 cores 32GB 200GB SSD
Medium 10-30 5-8 16 cores 64GB 500GB NVMe
Large 30-100 10-20 32 cores 128GB 1TB NVMe
XLarge 100+ 20+ 64 cores 256GB 2TB NVMe

2.3. System Reserved Resources

Kubernetes cần reserve resources cho system components trên mỗi node:

# /var/lib/kubelet/config.yaml
apiVersion: kubelet.config.k8s.io/v1beta1
kind: KubeletConfiguration
systemReserved:
  cpu: "500m"
  memory: "1Gi"
  ephemeral-storage: "10Gi"
kubeReserved:
  cpu: "500m"
  memory: "1Gi"
  ephemeral-storage: "5Gi"
evictionHard:
  memory.available: "500Mi"
  nodefs.available: "10%"
  imagefs.available: "15%"

Allocatable = Total - systemReserved - kubeReserved - evictionThreshold

Ví dụ: Worker node 16 cores, 64GB RAM:
  CPU allocatable:  16000m - 500m - 500m = 15000m
  RAM allocatable:  64Gi - 1Gi - 1Gi - 500Mi = 61.5Gi

PHẦN 3: SIZING STORAGE NODES (CEPH)

3.1. Ceph Components trên Storage Nodes


Storage Node
├── Ceph OSD daemon (1 per disk)    ─── Object Storage Daemon
│   ├── BlueStore (direct disk I/O)
│   └── WAL + DB on SSD (nếu dùng HDD)
├── Ceph MON (trên 3 nodes)         ─── Cluster monitor
├── Ceph MGR (trên 2 nodes)         ─── Manager, dashboard
└── kubelet + containerd + Cilium   ─── K8s agent

3.2. Tính toán Storage Capacity


Usable Capacity = Raw Capacity / Replication Factor × Utilization Target

Ví dụ: 3 nodes × 4 disks × 2TB = 24TB raw Replication factor = 3 (data replicated 3 lần) Utilization target = 75% (để headroom cho recovery)

Usable = 24TB / 3 × 0.75 = 6TB usable

Phân bổ: PostgreSQL data: 500GB (× 3 replicas nguồn PG) Kafka log retention: 500GB Loki logs: 1TB Thanos metrics: 500GB Velero backups: 1TB Application data: 500GB Buffer: 2TB ───────────────────────────── TỔNG: ~6TB ❯ Khớp 6TB usable

3.3. Sizing Ceph OSD Nodes

Component Sizing Rule Ví dụ (4 OSDs/node)
CPU per OSD 1 core per OSD (min) 4 cores cho OSDs
RAM per OSD 5GB per OSD (BlueStore default) 20GB cho OSDs
Ceph MON RAM ~2-4GB 4GB
System + K8s ~4GB RAM, 2 cores 4GB, 2 cores
TỔNG per node 6 cores, 28GB RAM

💡 Recommendation:


Storage Node (dedicated hoặc converged với worker):
  CPU:  8 cores
  RAM:  32-64GB (phụ thuộc số OSD)
  Disk: 1× 500GB NVMe (OS)
        4× 2TB NVMe (Ceph OSD)
  NIC:  2× 25GbE (1 cluster + 1 public network)

⚠️ Quyết định quan trọng: Dedicated storage nodes vs Converged (worker + storage)?


Dedicated Storage Nodes:
  ✅ Isolation: Storage I/O không ảnh hưởng workloads
  ✅ Independent scaling
  ❌ Thêm servers

Converged (Worker + Storage cùng node): ✅ Ít servers, tận dụng hardware ❌ Noisy neighbor: Ceph I/O có thể ảnh hưởng pods ❌ Node failure mất cả compute + storage

→ Production: Dedicated storage nodes → Lab/Small: Converged OK


PHẦN 4: NETWORK TOPOLOGY DESIGN

4.1. 4 Networks cho Production


┌─────────────────────────────────────────────────────────────────┐
│                    NETWORK TOPOLOGY                             │
├─────────────────────────────────────────────────────────────────┤
│                                                                 │
│  ┌── VLAN 10: Management Network (192.168.10.0/24) ──────────┐│
│  │  SSH access, monitoring, IPMI/iDRAC/iLO                    ││
│  │  MTU: 1500                                                  ││
│  │  NIC: eth0 (hoặc bond0 member)                             ││
│  └────────────────────────────────────────────────────────────┘│
│                                                                 │
│  ┌── VLAN 20: Cluster Network (10.10.20.0/24) ───────────────┐│
│  │  K8s API, Pod-to-Pod traffic, Service communication        ││
│  │  MTU: 9000 (Jumbo Frames)                                   ││
│  │  NIC: bond0 (eth1 + eth2 LACP)                             ││
│  └────────────────────────────────────────────────────────────┘│
│                                                                 │
│  ┌── VLAN 30: Storage Network (10.10.30.0/24) ───────────────┐│
│  │  Ceph cluster traffic (OSD replication, recovery)          ││
│  │  MTU: 9000 (Jumbo Frames)                                   ││
│  │  NIC: bond1 (eth3 + eth4 LACP) - Dedicated 25GbE          ││
│  └────────────────────────────────────────────────────────────┘│
│                                                                 │
│  ┌── VLAN 40: External Network (10.10.40.0/24) ──────────────┐│
│  │  User traffic, Ingress, MetalLB VIPs                       ││
│  │  MTU: 1500                                                  ││
│  │  NIC: bond0 (shared với Cluster, VLAN tagged)              ││
│  └────────────────────────────────────────────────────────────┘│
│                                                                 │
│  Firewall/Router: Giữa External ↔ Internal networks           │
│  DNS: Internal DNS cho *.k8s.local                             │
└─────────────────────────────────────────────────────────────────┘

4.2. IP Planning chi tiết

Node Management (VLAN 10) Cluster (VLAN 20) Storage (VLAN 30) Role
lb1 192.168.10.9 10.10.20.9 - HAProxy/keepalived primary
lb2 192.168.10.10 10.10.20.10 - HAProxy/keepalived backup
VIP - 10.10.20.100 - K8s API Server VIP
master1 192.168.10.11 10.10.20.11 - Control Plane 1
master2 192.168.10.12 10.10.20.12 - Control Plane 2
master3 192.168.10.13 10.10.20.13 - Control Plane 3
worker1 192.168.10.21 10.10.20.21 - Worker Node 1
worker2 192.168.10.22 10.10.20.22 - Worker Node 2
worker3 192.168.10.23 10.10.20.23 - Worker Node 3
storage1 192.168.10.31 10.10.20.31 10.10.30.31 Ceph OSD Node 1
storage2 192.168.10.32 10.10.20.32 10.10.30.32 Ceph OSD Node 2
storage3 192.168.10.33 10.10.20.33 10.10.30.33 Ceph OSD Node 3
MetalLB Pool - - - 10.10.40.200-250

4.3. NIC Bonding Configuration

NIC bonding (LACP) cung cấp link redundancy và bandwidth aggregation:

# /etc/netplan/01-bonding.yaml (Ubuntu 24.04)
network:
  version: 2
  renderer: networkd
  
  ethernets:
    eth0:
      dhcp4: false
    eth1:
      dhcp4: false
    eth2:
      dhcp4: false
    eth3:
      dhcp4: false
    eth4:
      dhcp4: false

  bonds:
    bond0:
      interfaces: [eth1, eth2]
      parameters:
        mode: 802.3ad         # LACP
        lacp-rate: fast
        mii-monitor-interval: 100
        transmit-hash-policy: layer3+4
      mtu: 9000

    bond1:
      interfaces: [eth3, eth4]
      parameters:
        mode: 802.3ad
        lacp-rate: fast
        mii-monitor-interval: 100
        transmit-hash-policy: layer3+4
      mtu: 9000

  vlans:
    bond0.20:
      id: 20
      link: bond0
      mtu: 9000
      addresses:
        - 10.10.20.21/24
      routes:
        - to: 10.244.0.0/16     # Pod CIDR
          via: 10.10.20.1
        - to: 10.96.0.0/12      # Service CIDR
          via: 10.10.20.1

    bond0.40:
      id: 40
      link: bond0
      addresses:
        - 10.10.40.21/24
      routes:
        - to: default
          via: 10.10.40.1

    bond1.30:
      id: 30
      link: bond1
      mtu: 9000
      addresses:
        - 10.10.30.21/24
# Apply cấu hình
sudo netplan apply

# Verify bonding
cat /proc/net/bonding/bond0
# Output:
# Bonding Mode: IEEE 802.3ad Dynamic link aggregation
# MII Status: up
# Slave Interface: eth1 → MII Status: up
# Slave Interface: eth2 → MII Status: up

# Verify VLAN
ip -d link show bond0.20

# Test MTU
ping -M do -s 8972 10.10.20.11  # 8972 + 28 = 9000 MTU

4.4. MTU Sizing

Network MTU Lý do
Management 1500 Standard, tương thích mọi device
Cluster (K8s) 9000 Jumbo frames giảm CPU overhead, tăng throughput
Storage (Ceph) 9000 Critical cho Ceph OSD replication performance
External 1500 Standard cho internet-facing traffic
Pod Network (Cilium) 8950 MTU underlay (9000) - VXLAN overhead (50)

⚠️ Jumbo Frames requirement: Tất cả switch ports trên path phải support và enable MTU 9000. Kiểm tra với switch admin trước khi deploy.


PHẦN 5: SWITCH VÀ FIREWALL REQUIREMENTS

5.1. Switch Requirements


Top-of-Rack (ToR) Switch Requirements:
  ├── L2/L3 capable
  ├── VLAN support (802.1Q)
  ├── LACP support (802.3ad)
  ├── Jumbo frames (MTU 9000)
  ├── Spanning Tree (RSTP/MSTP)
  └── Port count: 24-48 × 10/25GbE + 4-8 uplinks

Recommended Models (by budget): Budget: Arista 7010T, Dell S3048-ON Mid-range: Arista 7050SX, Cisco Nexus 93180YC Enterprise: Arista 7280R, Cisco Nexus 9336C

5.2. Firewall Rules giữa Networks

Source Destination Port Protocol Purpose
Management All nodes 22 TCP SSH
External Worker/LB 80, 443 TCP HTTP/HTTPS Ingress
External Master VIP 6443 TCP K8s API (nếu cần external)
Cluster Cluster 6443 TCP K8s API Server
Cluster Cluster 2379-2380 TCP etcd peer & client
Cluster Cluster 10250 TCP kubelet API
Cluster Cluster 10259 TCP kube-scheduler
Cluster Cluster 10257 TCP kube-controller-manager
Cluster Cluster 30000-32767 TCP NodePort range
Cluster Cluster 4240, 4244 TCP Cilium health, Hubble
Cluster Cluster 8472 UDP Cilium VXLAN
Storage Storage 6789 TCP Ceph MON
Storage Storage 6800-7300 TCP Ceph OSD
Cluster Storage 6789,6800-7300 TCP Ceph client access

PHẦN 6: DISK LAYOUT VÀ PARTITIONING

6.1. Control Plane Disk Layout


Disk: 1× 500GB NVMe SSD
├── /boot/efi     200MB   (EFI System Partition)
├── /boot         1GB     (kernel, initramfs)
├── /             50GB    (OS root)
├── /var/lib/etcd 100GB   (etcd data - SEPARATE partition!)
├── /var/lib/containerd 100GB (container images, layers)
├── /var/log      50GB    (system logs)
└── (remaining)   ~200GB  (buffer)

LVM recommended cho flexibility: VG: vg-system → LV: lv-root, lv-etcd, lv-containerd, lv-log

6.2. Worker Node Disk Layout


Disk: 1× 500GB NVMe SSD (OS)
├── /boot/efi     200MB
├── /boot         1GB
├── /             50GB
├── /var/lib/containerd 200GB (container images!)
├── /var/log      50GB
└── (remaining)   ~200GB
  • Raw disks cho Ceph OSD (nếu converged mode): /dev/sdb → Ceph OSD 0 /dev/sdc → Ceph OSD 1 (KHÔNG partition, KHÔNG format — Rook sẽ manage)

6.3. Storage Node Disk Layout


Disk 1: 500GB NVMe (OS)
├── / (OS root)
├── /var/lib/containerd
└── /var/log

Disk 2-5: 4× 2TB NVMe (Ceph OSD) /dev/nvme1n1 → Ceph OSD 0 (RAW - không format) /dev/nvme2n1 → Ceph OSD 1 (RAW) /dev/nvme3n1 → Ceph OSD 2 (RAW) /dev/nvme4n1 → Ceph OSD 3 (RAW)

⚠️ QUAN TRỌNG: KHÔNG tạo partition hay filesystem trên Ceph disks. Rook-Ceph sẽ sử dụng trực tiếp raw block devices.


PHẦN 7: BILL OF MATERIALS (BOM)

7.1. BOM cho Production (Medium - 20 microservices)

# Component Qty Specs Role
1 Server (Control Plane) 3 8C/32GB/500GB NVMe, 2×10GbE K8s masters
2 Server (Worker) 5 16C/64GB/500GB NVMe, 2×25GbE Workload nodes
3 Server (Storage) 3 8C/64GB/500GB NVMe + 4×2TB NVMe, 2×25GbE Ceph OSD
4 Server (LB) 2 4C/8GB/100GB SSD, 2×10GbE HAProxy/keepalived
5 ToR Switch 2 48×25GbE + 8×100GbE uplink Network
6 UPS 2 3kVA online double-conversion Power protection
7 PDU 2 Managed, dual feed Power distribution

Tổng: 13 servers + 2 switches + power infrastructure


💡 KEY TAKEAWAYS

  1. Control Plane cần NVMe SSD cho etcd, 8+ cores, 16-32GB RAM mỗi node
  2. Worker sizing tính từ tổng pod requests + 30% buffer + system reserved
  3. Ceph storage cần 5GB RAM per OSD, raw disks không format
  4. 4 networks riêng biệt bằng VLAN: Management, Cluster, Storage, External
  5. NIC bonding (LACP) cho link redundancy, Jumbo Frames 9000 cho cluster/storage
  6. Disk layout: etcd cần partition riêng, Ceph cần raw block devices

🎯 BÀI TẬP

Bài tập 1: Tính toán Sizing

Cho hệ thống gồm:

  • 30 microservices, mỗi service 2 replicas, avg 1 CPU/2GB RAM
  • PostgreSQL 3 nodes (4 CPU/8GB RAM each)
  • Kafka 3 brokers (4 CPU/8GB RAM each)
  • Full observability stack
  • Data retention: 90 ngày logs, 1 năm metrics

Tính: Số worker nodes, storage nodes, tổng disk capacity cần thiết.

Bài tập 2: Network Design

Vẽ network topology diagram chi tiết cho hệ thống trên, bao gồm:

  • VLAN assignment cho từng network segment
  • IP planning table cho tất cả nodes
  • NIC bonding topology
  • Firewall rules matrix

Bài tập 3: Lab Setup

  • Sử dụng Vagrantfile từ Bài 1, thêm NIC cho storage network
  • Cấu hình VLAN trên VMs (nếu hypervisor support)
  • Test connectivity: ping, iperf3 between all networks
  • Verify MTU 9000 hoạt động trên cluster network

📚 BÀI TIẾP THEO

Trong Bài 3: Chuẩn bị Linux OS và System Tuning, chúng ta sẽ cấu hình kernel parameters, tắt swap, setup NTP, SSH hardening và chuẩn bị tất cả nodes trước khi cài Kubernetes.