Chuyển đến nội dung chính

LESSON 2: PLANNING HARDWARE AND NETWORK TOPOLOGY

Calculate CPU/RAM/Disk for control plane, worker nodes, storage nodes. Network topology design: management network, cluster network, storage network, external network. VLAN, bonding, MTU sizing for production.

🔒 DevSecOps — Lesson 2 LESSON 2: PLANNING HARDWARE AND NETWORK TOPOLOGY

Deploy Microservices On-Premises with Kubernetes HA

Part 1: Platform & On-Premises Infrastructure Design

xdev.asia

🎯 LESSON OBJECTIVE__HTMLTAG_68___

After completing this lesson, you will:

  • ✅ Calculate accurate sizing for each type of node (Control Plane, Worker, Storage)
  • ✅ Design a production-grade network topology with VLAN separation__HTMLTAG_75___
  • ✅ Configure NIC bonding for HA networking
  • ✅ Understand and choose the appropriate MTU for each network segment
  • ✅ Create a complete hardware plan for a real project

PART 1: SIZING CONTROL PLANE NODES

1.1. Components running on Control Plane


Control Plane Node
├── kube-apiserver           ─── API endpoint, xử lý tất cả requests
├── etcd                     ─── Distributed KV store (cluster state)
├── kube-scheduler           ─── Pod scheduling decisions
├── kube-controller-manager  ─── Reconciliation loops
├── cloud-controller-manager ─── (Không dùng cho on-prem)
├── kubelet                  ─── Node agent
├── containerd               ─── Container runtime
└── Cilium agent            ─── CNI networking

1.2. Calculate Resources for Control Plane

etcd is the most critical element__HTMLTAG_91___

etcd performance depends mainly on disk I/O. Here are the sizing guidelines from etcd documentation:

Cluster Size Nodes Pods etcd CPU etcd RAM etcd Disk Disk Type
Small < 10 < 500 2 cores 4GB 50GB SSD
Medium 10-50 500-5000 4 cores 8GB 100GB NVMe SSD
Large 50-100 5000+ 8 cores 16GB 200GB NVMe SSD

⚠️ Critical: etcd requires disk latency p99 < 10ms. Dùng NVMe SSD chuyên dụng cho etcd.

kube-apiserver sizing


# Tính toán dựa trên số lượng requests/giây
API Server resources = f(number_of_nodes, number_of_pods, number_of_controllers)

Baseline (10 nodes, 500 pods): CPU: 2 cores RAM: 4GB

Scaling rule: +1 CPU per 1000 pods +2GB RAM per 1000 pods +1 CPU per 20 nodes

Synthesis of Control Plane Node Sizing

Component CPU Request CPU Limit RAM Request RAM Limit
kube-apiserver__HTMLTAG_193___ 250m 2000m 512Mi 4Gi
etcd 500m 4000m 1Gi 8Gi
kube-scheduler 100m 500m 128Mi 512Mi
kube-controller-manager 200m 1000m 256Mi 1Gi
kubelet + containerd 200m 500m 256Mi 1Gi
Cilium agent 100m 500m 256Mi 1Gi
OS overhead 500m - 1Gi -
TOTAL ~2 cores_ ~8 cores ~3.5Gi ~16Gi

💡 Recommendation for Production:


Control Plane Node (mỗi node):
  CPU:  8 cores (headroom cho burst)
  RAM:  16GB (minimum) - 32GB (recommended)
  Disk: 100GB NVMe SSD (OS + etcd)
        → etcd nên trên partition/disk riêng nếu có thể
  NIC:  2× 10GbE (bonding) hoặc 1× 25GbE

PART 2: SIZING WORKER NODES

2.1. Calculated from Workload Requirements

Formula to calculate Worker nodes:


Total Worker Resources = Σ (all pod requests) + System Reserved + Buffer

Ví dụ với hệ thống 20 microservices: ┌──────────────────────────────────────────────────────────────┐ │ Microservices (20 services × 2 replicas × avg 500m/1Gi): │ │ CPU: 20 × 2 × 500m = 20,000m = 20 cores │ │ RAM: 20 × 2 × 1Gi = 40Gi │ │ │ │ Databases (PostgreSQL 3 nodes, Redis 3, RabbitMQ 3, Kafka 3):│ │ CPU: 12 × 2000m = 24,000m = 24 cores │ │ RAM: 12 × 4Gi = 48Gi │ │ │ │ Observability (Prometheus×2, Grafana, Loki, Tempo, Alloy): │ │ CPU: ~8 cores │ │ RAM: ~24Gi │ │ │ │ Platform (ArgoCD, Vault, Istio, cert-manager, Kyverno): │ │ CPU: ~6 cores │ │ RAM: ~16Gi │ │ │ │ System Reserved per node (kubelet, containerd, Cilium, OS): │ │ CPU: ~1.5 cores × N nodes │ │ RAM: ~2Gi × N nodes │ │ │ │ TỔNG REQUEST: │ │ CPU: ~58 cores + (1.5 × N) │ │ RAM: ~128Gi + (2 × N) │ └──────────────────────────────────────────────────────────────┘

Buffer (30% headroom cho HA + burst): CPU: 58 × 1.3 = ~76 cores RAM: 128 × 1.3 = ~167Gi

Sizing calculation: Nếu mỗi worker: 16 cores, 64GB RAM Số workers = max(76/16, 167/64) = max(4.75, 2.6) = 5 workers

💡 Khuyến nghị: 5-6 worker nodes × (16 cores, 64GB RAM) → Cho phép mất 1 node mà workloads vẫn schedulable

2.2. Sizing Guidelines by Size

Scale Services Workers CPU/node RAM/node Disk/node
Small (Lab) 5-10 3 8 cores 32GB 200GB SSD
Medium 10-30 5-8 16 cores 64GB 500GB NVMe
Large 30-100 10-20 32 cores 128GB 1TB NVMe
XLarge 100+ 20+ 64 cores 256GB 2TB NVMe

2.3. System Reserved Resources

Kubernetes needs to reserve resources for system components on each node:

# /var/lib/kubelet/config.yaml
apiVersion: kubelet.config.k8s.io/v1beta1
kind: KubeletConfiguration
systemReserved:
  cpu: "500m"
  memory: "1Gi"
  ephemeral-storage: "10Gi"
kubeReserved:
  cpu: "500m"
  memory: "1Gi"
  ephemeral-storage: "5Gi"
evictionHard:
  memory.available: "500Mi"
  nodefs.available: "10%"
  imagefs.available: "15%"

Allocatable = Total - systemReserved - kubeReserved - evictionThreshold

Ví dụ: Worker node 16 cores, 64GB RAM:
  CPU allocatable:  16000m - 500m - 500m = 15000m
  RAM allocatable:  64Gi - 1Gi - 1Gi - 500Mi = 61.5Gi

PART 3: SIZING STORAGE NODES (CEPH)

3.1. Ceph Components on Storage Nodes


Storage Node
├── Ceph OSD daemon (1 per disk)    ─── Object Storage Daemon
│   ├── BlueStore (direct disk I/O)
│   └── WAL + DB on SSD (nếu dùng HDD)
├── Ceph MON (trên 3 nodes)         ─── Cluster monitor
├── Ceph MGR (trên 2 nodes)         ─── Manager, dashboard
└── kubelet + containerd + Cilium   ─── K8s agent

3.2. Calculate Storage Capacity


Usable Capacity = Raw Capacity / Replication Factor × Utilization Target

Ví dụ: 3 nodes × 4 disks × 2TB = 24TB raw Replication factor = 3 (data replicated 3 lần) Utilization target = 75% (để headroom cho recovery)

Usable = 24TB / 3 × 0.75 = 6TB usable

Phân bổ: PostgreSQL data: 500GB (× 3 replicas nguồn PG) Kafka log retention: 500GB Loki logs: 1TB Thanos metrics: 500GB Velero backups: 1TB Application data: 500GB Buffer: 2TB ───────────────────────────── TỔNG: ~6TB ❯ Khớp 6TB usable

3.3. Sizing Ceph OSD Nodes

Component Sizing Rule Example (4 OSDs/node)
CPU per OSD 1 core per OSD (min) 4 cores for OSDs
RAM per OSD 5GB per OSD (BlueStore default) 20GB for OSDs
Ceph MON RAM ~2-4GB 4GB
System + K8s ~4GB RAM, 2 cores 4GB, 2 cores
TOTAL per node 6 cores, 28GB RAM

💡 Recommendation:


Storage Node (dedicated hoặc converged với worker):
  CPU:  8 cores
  RAM:  32-64GB (phụ thuộc số OSD)
  Disk: 1× 500GB NVMe (OS)
        4× 2TB NVMe (Ceph OSD)
  NIC:  2× 25GbE (1 cluster + 1 public network)

⚠️ Important decision: Dedicated storage nodes vs Converged (worker + storage)?


Dedicated Storage Nodes:
  ✅ Isolation: Storage I/O không ảnh hưởng workloads
  ✅ Independent scaling
  ❌ Thêm servers

Converged (Worker + Storage cùng node): ✅ Ít servers, tận dụng hardware ❌ Noisy neighbor: Ceph I/O có thể ảnh hưởng pods ❌ Node failure mất cả compute + storage

→ Production: Dedicated storage nodes → Lab/Small: Converged OK


PART 4: NETWORK TOPOLOGY DESIGN

4.1. 4 Networks for Production


┌─────────────────────────────────────────────────────────────────┐
│                    NETWORK TOPOLOGY                             │
├─────────────────────────────────────────────────────────────────┤
│                                                                 │
│  ┌── VLAN 10: Management Network (192.168.10.0/24) ──────────┐│
│  │  SSH access, monitoring, IPMI/iDRAC/iLO                    ││
│  │  MTU: 1500                                                  ││
│  │  NIC: eth0 (hoặc bond0 member)                             ││
│  └────────────────────────────────────────────────────────────┘│
│                                                                 │
│  ┌── VLAN 20: Cluster Network (10.10.20.0/24) ───────────────┐│
│  │  K8s API, Pod-to-Pod traffic, Service communication        ││
│  │  MTU: 9000 (Jumbo Frames)                                   ││
│  │  NIC: bond0 (eth1 + eth2 LACP)                             ││
│  └────────────────────────────────────────────────────────────┘│
│                                                                 │
│  ┌── VLAN 30: Storage Network (10.10.30.0/24) ───────────────┐│
│  │  Ceph cluster traffic (OSD replication, recovery)          ││
│  │  MTU: 9000 (Jumbo Frames)                                   ││
│  │  NIC: bond1 (eth3 + eth4 LACP) - Dedicated 25GbE          ││
│  └────────────────────────────────────────────────────────────┘│
│                                                                 │
│  ┌── VLAN 40: External Network (10.10.40.0/24) ──────────────┐│
│  │  User traffic, Ingress, MetalLB VIPs                       ││
│  │  MTU: 1500                                                  ││
│  │  NIC: bond0 (shared với Cluster, VLAN tagged)              ││
│  └────────────────────────────────────────────────────────────┘│
│                                                                 │
│  Firewall/Router: Giữa External ↔ Internal networks           │
│  DNS: Internal DNS cho *.k8s.local                             │
└─────────────────────────────────────────────────────────────────┘

4.2. IP Planning details

Node Management (VLAN 10) Cluster (VLAN 20) Storage (VLAN 30) Role
lb1 192.168.10.9 10.10.20.9 - HAProxy/keepalived primary
lb2 192.168.10.10 10.10.20.10 - HAProxy/keepalived backup
VIP - 10.10.20.100 - K8s API Server VIP
master1 192.168.10.11 10.10.20.11 - Control Plane 1
master2 192.168.10.12 10.10.20.12 - Control Plane 2
master3 192.168.10.13 10.10.20.13 - Control Plane 3
worker1 192.168.10.21 10.10.20.21 - Worker Node 1
worker2 192.168.10.22 10.10.20.22 - Worker Node 2
worker3 192.168.10.23 10.10.20.23 - Worker Node 3
storage1 192.168.10.31 10.10.20.31 10.10.30.31 Ceph OSD Node 1
storage2 192.168.10.32 10.10.20.32 10.10.30.32 Ceph OSD Node 2
storage3 192.168.10.33 10.10.20.33 10.10.30.33 Ceph OSD Node 3
MetalLB Pool - - - 10.10.40.200-250

4.3. NIC Bonding Configuration

NIC bonding (LACP) provides link redundancy and bandwidth aggregation:

# /etc/netplan/01-bonding.yaml (Ubuntu 24.04)
network:
  version: 2
  renderer: networkd
  
  ethernets:
    eth0:
      dhcp4: false
    eth1:
      dhcp4: false
    eth2:
      dhcp4: false
    eth3:
      dhcp4: false
    eth4:
      dhcp4: false

  bonds:
    bond0:
      interfaces: [eth1, eth2]
      parameters:
        mode: 802.3ad         # LACP
        lacp-rate: fast
        mii-monitor-interval: 100
        transmit-hash-policy: layer3+4
      mtu: 9000

    bond1:
      interfaces: [eth3, eth4]
      parameters:
        mode: 802.3ad
        lacp-rate: fast
        mii-monitor-interval: 100
        transmit-hash-policy: layer3+4
      mtu: 9000

  vlans:
    bond0.20:
      id: 20
      link: bond0
      mtu: 9000
      addresses:
        - 10.10.20.21/24
      routes:
        - to: 10.244.0.0/16     # Pod CIDR
          via: 10.10.20.1
        - to: 10.96.0.0/12      # Service CIDR
          via: 10.10.20.1

    bond0.40:
      id: 40
      link: bond0
      addresses:
        - 10.10.40.21/24
      routes:
        - to: default
          via: 10.10.40.1

    bond1.30:
      id: 30
      link: bond1
      mtu: 9000
      addresses:
        - 10.10.30.21/24
# Apply cấu hình
sudo netplan apply

# Verify bonding
cat /proc/net/bonding/bond0
# Output:
# Bonding Mode: IEEE 802.3ad Dynamic link aggregation
# MII Status: up
# Slave Interface: eth1 → MII Status: up
# Slave Interface: eth2 → MII Status: up

# Verify VLAN
ip -d link show bond0.20

# Test MTU
ping -M do -s 8972 10.10.20.11  # 8972 + 28 = 9000 MTU

4.4. MTU Sizing

Network MTU Reason
Management 1500 Standard, compatible with all devices
Cluster (K8s) 9000 Jumbo frames reduce CPU overhead, increase throughput__HTMLTAG_688___
Storage (Ceph) 9000 Critical for Ceph OSD replication performance__HTMLTAG_696___
External 1500 Standard for internet-facing traffic__HTMLTAG_704___
Pod Network (Cilium) 8950 MTU underlay (9000) - VXLAN overhead (50)

⚠️ Jumbo Frames requirement: All switch ports on the path must support and enable MTU 9000. Check with switch admin before deploying.


PART 5: SWITCH AND FIREWALL REQUIREMENTS

5.1. Switch Requirements


Top-of-Rack (ToR) Switch Requirements:
  ├── L2/L3 capable
  ├── VLAN support (802.1Q)
  ├── LACP support (802.3ad)
  ├── Jumbo frames (MTU 9000)
  ├── Spanning Tree (RSTP/MSTP)
  └── Port count: 24-48 × 10/25GbE + 4-8 uplinks

Recommended Models (by budget): Budget: Arista 7010T, Dell S3048-ON Mid-range: Arista 7050SX, Cisco Nexus 93180YC Enterprise: Arista 7280R, Cisco Nexus 9336C

5.2. Firewall Rules between Networks

Source Destination Port Protocol Purpose
Management All nodes 22 TCP SSH
External Worker/LB 80, 443 TCP HTTP/HTTPS Ingress
External Master VIP 6443 TCP K8s API (if needed external)
Cluster Cluster 6443 TCP K8s API Server
Cluster Cluster 2379-2380 TCP etcd peer & client
Cluster Cluster 10250 TCP kubelet API
Cluster Cluster 10259 TCP kube-scheduler__HTMLTAG_827___
Cluster Cluster 10257 TCP kube-controller-manager
Cluster Cluster 30000-32767 TCP NodePort range
Cluster Cluster 4240, 4244 TCP Cilium health, Hubble
Cluster Cluster 8472 UDP Cilium VXLAN
Storage Storage 6789 TCP Ceph MO_
Storage Storage 6800-7300 TCP Ceph OSD
Cluster Storage 6789,6800-7300 TCP Ceph client access

PART 6: DISK LAYOUT AND PARTITIONING__HTMLTAG_918___

6.1. Control Plane Disk Layout


Disk: 1× 500GB NVMe SSD
├── /boot/efi     200MB   (EFI System Partition)
├── /boot         1GB     (kernel, initramfs)
├── /             50GB    (OS root)
├── /var/lib/etcd 100GB   (etcd data - SEPARATE partition!)
├── /var/lib/containerd 100GB (container images, layers)
├── /var/log      50GB    (system logs)
└── (remaining)   ~200GB  (buffer)

LVM recommended cho flexibility: VG: vg-system → LV: lv-root, lv-etcd, lv-containerd, lv-log

6.2. Worker Node Disk Layout


Disk: 1× 500GB NVMe SSD (OS)
├── /boot/efi     200MB
├── /boot         1GB
├── /             50GB
├── /var/lib/containerd 200GB (container images!)
├── /var/log      50GB
└── (remaining)   ~200GB
  • Raw disks cho Ceph OSD (nếu converged mode): /dev/sdb → Ceph OSD 0 /dev/sdc → Ceph OSD 1 (KHÔNG partition, KHÔNG format — Rook sẽ manage)

6.3. Storage Node Disk Layout


Disk 1: 500GB NVMe (OS)
├── / (OS root)
├── /var/lib/containerd
└── /var/log

Disk 2-5: 4× 2TB NVMe (Ceph OSD) /dev/nvme1n1 → Ceph OSD 0 (RAW - không format) /dev/nvme2n1 → Ceph OSD 1 (RAW) /dev/nvme3n1 → Ceph OSD 2 (RAW) /dev/nvme4n1 → Ceph OSD 3 (RAW)

⚠️ QUAN TRỌNG: KHÔNG tạo partition hay filesystem trên Ceph disks. Rook-Ceph sẽ sử dụng trực tiếp raw block devices.


PART 7: BILL OF MATERIALS (BOM)

7.1. BOM for Production (Medium - 20 microservices)

# Component Qty Specs Role
1 Server (Control Plane) 3 8C/32GB/500GB NVMe, 2×10GbE K8s masters
2 Server (Worker) 5 16C/64GB/500GB NVMe, 2×25GbE Workload nodes
3 Server (Storage) 3 8C/64GB/500GB NVMe + 4×2TB NVMe, 2×25GbE Ceph OSD
4 Server (LB) 2 4C/8GB/100GB SSD, 2×10GbE HAProxy/keepalived
5 ToR Switch 2 48×25GbE + 8×100GbE uplink Network
6 UPS 2 3kVA online double-conversion Power protection
7 PDU 2 Managed, dual feed Power distribution

Total: 13 servers + 2 switches + power infrastructure


💡 KEY TAKEAWAYS

  1. Control Plane needs NVMe SSD for etcd, 8+ cores, 16-32GB RAM per node
  2. Worker sizing calculated from total pod requests + 30% buffer + system reserved
  3. Ceph storageneeds 5GB RAM per OSD, raw disks not formatted
  4. 4 separate networks by VLAN: Management, Cluster, Storage, External
  5. NIC bonding (LACP) for link redundancy, Jumbo Frames 9000 for cluster/storage
  6. Disk layout: etcd needs separate partition, Ceph needs raw block devices

🎯 EXERCISES__HTMLTAG_1069___

Exercise 1: Calculating Sizing

For systems including:

  • 30 microservices, each service 2 replicas, avg 1 CPU/2GB RAM
  • PostgreSQL 3 nodes (4 CPU/8GB RAM each)
  • Kafka 3 brokers (4 CPU/8GB RAM each)
  • Full observability stack
  • Data retention: 90 days logs, 1 year metrics

Calculate: Number of worker nodes, storage nodes, total disk capacity needed.

Exercise 2: Network Design__HTMLTAG_1089___

Draw a detailed network topology diagram for the above system, including:

  • VLAN assignment for each network segment__HTMLTAG_1094___
  • IP planning table for all nodes__HTMLTAG_1096___
  • NIC bonding topology
  • Firewall rules matrix

Exercise 3: Lab Setup__HTMLTAG_1103___
  • Using Vagrantfile from Lesson 1, add NIC for storage network
  • Configure VLANs on VMs (if hypervisor supports)
  • Test connectivity: ping, iperf3 between all networks
  • Verify MTU 9000 works on cluster network

📚 NEXT POST

In Lesson 3: Preparing Linux OS and System Tuning, we will configure kernel parameters, turn off swap, setup NTP, SSH hardening and prepare all nodes before installing Kubernetes.