Chuyển đến nội dung chính

LESSON 3: PREPARING LINUX OS AND SYSTEM TUNING

Install Ubuntu 24.04/RHEL 9, configure kernel parameters for K8s (net.bridge, ip_forward, inotify), turn off swap, configure chrony/NTP, firewall rules, SSH hardening and prepare all nodes before installing K8s.

🔒 DevSecOps — Lesson 3 LESSON 3: PREPARING LINUX OS AND SYSTEM TUNING

Deploy Microservices On-Premises with Kubernetes HA

Part 1: Platform & On-Premises Infrastructure Design

xdev.asia

🎯 LESSON OBJECTIVE__HTMLTAG_66___

After completing this lesson, you will:

  • ✅ Install and update Ubuntu 24.04 LTS for all nodes
  • ✅ Configure required kernel parameters for Kubernetes
  • ✅ Turn off swap permanently and understand why
  • ✅ Set up NTP synchronization for the entire cluster__HTMLTAG_77___
  • ✅ Hardening SSH and configuring firewall rules
  • ✅ Standardize hostname, hosts file and DNS resolution

PART 1: INSTALL UBUNTU 24.04 LTS

1.1. Why Ubuntu 24.04 LTS?

  • LTS (Long Term Support): 12 years support (until 2036 with Ubuntu Pro)
  • Kernel 6.8+: Good support for eBPF (Cilium), cgroup v2, io_uring
  • systemd 255+: Improve cgroup v2 management
  • Widespread adoption: Thoroughly tested with K8s, Ceph, and CNCF tools

⚠️ Note: RHEL 9 / Rocky Linux 9 is also a good choice for enterprises. This guide uses Ubuntu but will note RHEL commands when different.

1.2. Basic settings

# Ubuntu Server 24.04 minimal installation
# Chọn options:
#   - Minimal server (không GUI)
#   - OpenSSH server
#   - Disk layout: LVM (theo Bài 2)
#   - Timezone: Asia/Ho_Chi_Minh (hoặc UTC)

Sau khi cài xong, update hệ thống

sudo apt update && sudo apt upgrade -y

Cài đặt essential packages

sudo apt install -y
curl
wget
gnupg
apt-transport-https
ca-certificates
software-properties-common
net-tools
ipvsadm
ipset
jq
bash-completion
vim
htop
iotop
sysstat
lsof
tcpdump
conntrack
socat
nfs-common
open-iscsi
lvm2
tree

RHEL 9 equivalent:

sudo dnf update -y

sudo dnf install -y curl wget ...


PART 2: HOSTNAME AND DNS RESOLUTION__HTMLTAG_114___

2.1. Set hostname on each node

# Trên mỗi node, đặt hostname tương ứng
# master1:
sudo hostnamectl set-hostname master1

master2:

sudo hostnamectl set-hostname master2

master3:

sudo hostnamectl set-hostname master3

worker1-3, storage1-3, lb1-2 tương tự

2.2. Configure /etc/hosts on ALL nodes

cat >> /etc/hosts << 'EOF'

K8s HA Cluster - Management Network

192.168.10.9 lb1 192.168.10.10 lb2 192.168.10.11 master1 192.168.10.12 master2 192.168.10.13 master3 192.168.10.21 worker1 192.168.10.22 worker2 192.168.10.23 worker3 192.168.10.31 storage1 192.168.10.32 storage2 192.168.10.33 storage3

K8s API Server VIP (Cluster Network)

10.10.20.100 k8s-api.local

EOF

💡 Tip: In production, use Internal DNS server (CoreDNS or BIND) instead of /etc/hosts. This lesson uses /etc/hosts for simplicity.


PART 3: KERNEL PARAMETERS FOR KUBERNETES

3.1. Why is it necessary to customize the kernel?

Kubernetes requires some kernel features to be enabled:

Parameter Value Reason
net.bridge.bridge-nf-call-iptables__HTMLTAG_145___ 1 Bridge traffic via iptables rules (required for Service)
net.bridge.bridge-nf-call-ip6tables 1 IPv6 bridge traffic via ip6tables__HTMLTAG_157___
net.ipv4.ip_forward 1 Forward packets between interfaces (pod networking)
net.ipv6.conf.all.forwarding 1 IPv6 forwarding (if using dual-stack)
fs.inotify.max_user_instances 8192 For many containers need inotify (file watching)
fs.inotify.max_user_watches 524288 Inotify watches per user

3.2. Configure Kernel Parameters

# Load required kernel modules
cat > /etc/modules-load.d/k8s.conf << 'EOF'
overlay
br_netfilter
ip_vs
ip_vs_rr
ip_vs_wrr
ip_vs_sh
nf_conntrack
EOF

Load modules ngay lập tức

modprobe overlay modprobe br_netfilter modprobe ip_vs modprobe ip_vs_rr modprobe ip_vs_wrr modprobe ip_vs_sh modprobe nf_conntrack

Verify modules loaded

lsmod | grep -E "overlay|br_netfilter|ip_vs"

# Cấu hình sysctl parameters
cat > /etc/sysctl.d/99-kubernetes.conf << 'EOF'
# ─── Required cho Kubernetes ───
net.bridge.bridge-nf-call-iptables  = 1
net.bridge.bridge-nf-call-ip6tables = 1
net.ipv4.ip_forward                 = 1
net.ipv6.conf.all.forwarding        = 1

# ─── Performance tuning ───
# Tăng conntrack table cho nhiều connections
net.netfilter.nf_conntrack_max = 1048576

# Tăng số file descriptors
fs.file-max = 2097152
fs.nr_open  = 1048576

# inotify cho containers
fs.inotify.max_user_instances = 8192
fs.inotify.max_user_watches   = 524288

# ─── Network performance ───
# Tăng socket buffer sizes
net.core.somaxconn             = 65535
net.core.netdev_max_backlog    = 65535
net.ipv4.tcp_max_syn_backlog   = 65535

# TCP keepalive (cho long-lived connections)
net.ipv4.tcp_keepalive_time    = 600
net.ipv4.tcp_keepalive_intvl   = 30
net.ipv4.tcp_keepalive_probes  = 10

# Tăng port range
net.ipv4.ip_local_port_range   = 1024 65535

# ─── Memory management ───
vm.max_map_count     = 262144    # Required cho Elasticsearch/Kafka
vm.swappiness        = 0         # Minimize swap usage
vm.overcommit_memory = 1         # Allow memory overcommit (Redis)

# ─── Ceph OSD tuning (chỉ cần trên storage nodes) ───
# kernel.pid_max = 4194304
# vm.min_free_kbytes = 1048576
EOF

# Apply ngay lập tức
sysctl --system

# Verify
sysctl net.bridge.bridge-nf-call-iptables
# Output: net.bridge.bridge-nf-call-iptables = 1

sysctl net.ipv4.ip_forward
# Output: net.ipv4.ip_forward = 1

PART 4: TURN OFF SWAP

4.1. Why must Swap be turned off?

Kubelet requires swap to be disabled__HTMLTAG_203___ (although K8s 1.28+ has beta swap support). Reason:

  • Swap causes unpredictable latency for containers
  • K8s scheduler calculates resources based on physical memory
  • Swap hidden OOM issues, making debugging difficult
  • etcd and databases perform very poorly when swapping__HTMLTAG_219___

4.2. Turn off Swap permanently

# Tắt swap ngay lập tức
sudo swapoff -a

Tắt vĩnh viễn: comment dòng swap trong /etc/fstab

sudo sed -i '/ swap / s/^/#/' /etc/fstab

Verify

free -h | grep Swap

Output:

Swap: 0B 0B 0B

Double verify - không có swap partitions

cat /proc/swaps

Output: (empty)

Nếu dùng swap file:

sudo rm -f /swap.img # Xóa swap file nếu có


PART 5: NTP TIME SYNCHRONIZATION

5.1. Why is NTP extremely important?

  • etcd: Use timestamps for leader election, requires clock skew < 500ms
  • TLS certificates: Verify based on time, wrong clock → cert invalid
  • Logs correlation: Log timestamps must match between nodes
  • Ceph: MON quorum request clock skew < 50ms

5.2. Chrony configuration (NTP)

# Cài đặt chrony (thay thế ntpd, tiêu chuẩn cho RHEL/Ubuntu mới)
sudo apt install -y chrony

Cấu hình chrony

cat > /etc/chrony/chrony.conf << 'EOF'

NTP servers - dùng pool gần nhất

pool ntp.ubuntu.com iburst maxsources 4 pool time.google.com iburst maxsources 2 pool time.cloudflare.com iburst maxsources 2

Fallback: GPS hoặc internal NTP server

server ntp.internal.company.com iburst prefer

Record rate of clock drift

driftfile /var/lib/chrony/chrony.drift

Allow NTP client access from cluster network

allow 192.168.10.0/24 allow 10.10.20.0/24

Step clock nếu offset > 1 giây trong 3 updates đầu

makestep 1.0 3

Enable hardware timestamping nếu NIC support

hwtimestamp *

Enable kernel synchronization

rtcsync

Log

logdir /var/log/chrony EOF

Restart chrony

sudo systemctl restart chrony sudo systemctl enable chrony

Verify synchronization

chronyc sources -v

Output:

^* ntp.ubuntu.com 2 6 17 64 +0.153ms 0.421ms 0.312ms

chronyc tracking

Output:

Reference ID : A29FC801 (time.google.com)

Stratum : 2

System time : 0.000000023 seconds fast of NTP time

Kiểm tra trên tất cả nodes: clock skew < 1ms

for host in master{1..3} worker{1..3} storage{1..3}; do echo -n "$host: " ssh $host "chronyc tracking | grep 'System time'" done


PART 6: SSH HARDENING

6.1. Secure SSH configuration

# Backup cấu hình gốc
sudo cp /etc/ssh/sshd_config /etc/ssh/sshd_config.bak

Tạo cấu hình SSH hardened

cat > /etc/ssh/sshd_config.d/99-hardening.conf << 'EOF'

─── Authentication ───

PermitRootLogin prohibit-password # Chỉ cho SSH key, không password PasswordAuthentication no # Tắt password authentication PubkeyAuthentication yes # Chỉ dùng SSH keys AuthenticationMethods publickey # Enforce public key only

─── Security ───

MaxAuthTries 3 # Max 3 lần thử MaxSessions 10 # Max 10 sessions/connection LoginGraceTime 30 # 30s timeout cho login ClientAliveInterval 300 # Gửi keepalive mỗi 5 phút ClientAliveCountMax 3 # Disconnect sau 3 lần không reply PermitEmptyPasswords no X11Forwarding no # Không cần X11 AllowTcpForwarding yes # Cần cho kubectl port-forward AllowAgentForwarding yes

─── Ciphers (strong only) ───

Ciphers [email protected],[email protected],[email protected] MACs [email protected],[email protected] KexAlgorithms [email protected],[email protected]

─── Access control ───

AllowUsers root admin deploy # Chỉ cho phép users cụ thể

AllowGroups k8s-admins # Hoặc theo group

EOF

Validate config trước khi apply

sudo sshd -t

Output: (no errors)

Apply

sudo systemctl restart sshd

6.2. Setup SSH Key Authentication

# Trên workstation/jump host:
# Tạo dedicated key cho K8s cluster (đã làm ở Bài 1)
ssh-keygen -t ed25519 -C "[email protected]" -f ~/.ssh/k8s-admin

Copy sang tất cả nodes

for host in lb{1,2} master{1..3} worker{1..3} storage{1..3}; do ssh-copy-id -i ~/.ssh/k8s-admin.pub root@${host} done

Test kết nối không cần password

ssh -i ~/.ssh/k8s-admin root@master1 "hostname && date"


PART 7: FIREWALL CONFIGURATION

7.1. UFW for Control Plane Nodes

# Trên master1, master2, master3:

Reset firewall

sudo ufw --force reset sudo ufw default deny incoming sudo ufw default allow outgoing

SSH

sudo ufw allow from 192.168.10.0/24 to any port 22 proto tcp comment 'SSH from mgmt'

K8s API Server

sudo ufw allow 6443/tcp comment 'K8s API Server'

etcd

sudo ufw allow from 10.10.20.0/24 to any port 2379:2380 proto tcp comment 'etcd'

kubelet API

sudo ufw allow 10250/tcp comment 'kubelet API'

kube-scheduler, kube-controller-manager

sudo ufw allow 10259/tcp comment 'kube-scheduler' sudo ufw allow 10257/tcp comment 'kube-controller-manager'

Cilium

sudo ufw allow 4240/tcp comment 'Cilium health' sudo ufw allow 4244/tcp comment 'Hubble' sudo ufw allow 8472/udp comment 'Cilium VXLAN'

VRRP (keepalived) - nếu LB trên cùng node

sudo ufw allow proto vrrp from 10.10.20.0/24 comment 'keepalived VRRP'

Enable

sudo ufw enable sudo ufw status verbose

7.2. UFW for Worker Nodes

# Trên worker1, worker2, worker3:
sudo ufw --force reset
sudo ufw default deny incoming
sudo ufw default allow outgoing

SSH

sudo ufw allow from 192.168.10.0/24 to any port 22 proto tcp

kubelet API

sudo ufw allow 10250/tcp

NodePort range

sudo ufw allow 30000:32767/tcp comment 'K8s NodePort'

Cilium

sudo ufw allow 4240/tcp sudo ufw allow 4244/tcp sudo ufw allow 8472/udp

Ceph client (nếu converged mode)

sudo ufw allow from 10.10.20.0/24 to any port 6789 proto tcp comment 'Ceph MON' sudo ufw allow from 10.10.20.0/24 to any port 6800:7300 proto tcp comment 'Ceph OSD'

sudo ufw enable

💡 Alternative: Many production environments use nftables or iptables directly instead of UFW. Cilium can also replace node-level firewalls with Host Policies.


PART 8: OTHER CUSTOMIZATIONS

8.1. Disable Unattended Upgrades (Production)

# Trong production, tự kiểm soát upgrades
sudo apt remove -y unattended-upgrades
# Hoặc cấu hình chỉ security updates:
# sudo dpkg-reconfigure -plow unattended-upgrades

8.2. Configure ulimits

# Tăng limits cho tất cả users/processes
cat >> /etc/security/limits.conf << 'EOF'
*       soft    nofile      1048576
*       hard    nofile      1048576
*       soft    nproc       65535
*       hard    nproc       65535
*       soft    memlock     unlimited
*       hard    memlock     unlimited
EOF

Đảm bảo PAM đọc limits.conf

grep -q 'pam_limits.so' /etc/pam.d/common-session ||
echo "session required pam_limits.so" >> /etc/pam.d/common-session

8.3. Disable Transparent Huge Pages

# THP gây latency spikes cho databases (PostgreSQL, Redis)
cat > /etc/systemd/system/disable-thp.service << 'EOF'
[Unit]
Description=Disable Transparent Huge Pages
DefaultDependencies=no
After=sysinit.target local-fs.target
Before=basic.target

[Service] Type=oneshot ExecStart=/bin/sh -c 'echo never > /sys/kernel/mm/transparent_hugepage/enabled' ExecStart=/bin/sh -c 'echo never > /sys/kernel/mm/transparent_hugepage/defrag'

[Install] WantedBy=basic.target EOF

sudo systemctl daemon-reload sudo systemctl enable --now disable-thp.service

Verify

cat /sys/kernel/mm/transparent_hugepage/enabled

Output: always madvise [never]

8.4. cgroup v2 Verification

# Ubuntu 24.04 mặc định dùng cgroup v2
# Verify:
stat -fc %T /sys/fs/cgroup/
# Output: cgroup2fs  ← cgroup v2 ✅
# Nếu output là tmpfs → cgroup v1, cần migrate

Nếu cần force enable cgroup v2:

sudo sed -i 's/GRUB_CMDLINE_LINUX=""/GRUB_CMDLINE_LINUX="systemd.unified_cgroup_hierarchy=1"/' /etc/default/grub

sudo update-grub && sudo reboot


PART 9: AUTOMATION — SCRIPT PREPARING ALL NODES

9.1. Node preparation script (runs on all K8s nodes)

#!/bin/bash
# prepare-k8s-node.sh
# Chạy script này trên tất cả nodes (masters + workers)
set -euo pipefail

echo "=== [1/8] Updating system ===" apt update && apt upgrade -y

echo "=== [2/8] Installing packages ===" apt install -y curl wget gnupg apt-transport-https ca-certificates
software-properties-common net-tools ipvsadm ipset jq bash-completion
vim htop iotop sysstat lsof tcpdump conntrack socat nfs-common
open-iscsi lvm2 chrony tree

echo "=== [3/8] Loading kernel modules ===" cat > /etc/modules-load.d/k8s.conf << 'MODULES' overlay br_netfilter ip_vs ip_vs_rr ip_vs_wrr ip_vs_sh nf_conntrack MODULES

modprobe overlay modprobe br_netfilter modprobe ip_vs ip_vs_rr ip_vs_wrr ip_vs_sh modprobe nf_conntrack

echo "=== [4/8] Configuring sysctl ===" cat > /etc/sysctl.d/99-kubernetes.conf << 'SYSCTL' net.bridge.bridge-nf-call-iptables = 1 net.bridge.bridge-nf-call-ip6tables = 1 net.ipv4.ip_forward = 1 net.netfilter.nf_conntrack_max = 1048576 fs.file-max = 2097152 fs.inotify.max_user_instances = 8192 fs.inotify.max_user_watches = 524288 net.core.somaxconn = 65535 net.ipv4.tcp_keepalive_time = 600 net.ipv4.ip_local_port_range = 1024 65535 vm.max_map_count = 262144 vm.swappiness = 0 SYSCTL sysctl --system

echo "=== [5/8] Disabling swap ===" swapoff -a sed -i '/ swap / s/^/#/' /etc/fstab

echo "=== [6/8] Configuring chrony NTP ===" systemctl enable --now chrony

echo "=== [7/8] Disabling THP ===" echo never > /sys/kernel/mm/transparent_hugepage/enabled echo never > /sys/kernel/mm/transparent_hugepage/defrag

echo "=== [8/8] Verifying ===" echo "--- Swap status ---" free -h | grep Swap echo "--- cgroup version ---" stat -fc %T /sys/fs/cgroup/ echo "--- Kernel modules ---" lsmod | grep -E "overlay|br_netfilter|ip_vs" | awk '{print $1}' echo "--- Key sysctl values ---" sysctl net.bridge.bridge-nf-call-iptables net.ipv4.ip_forward vm.swappiness

echo "" echo "✅ Node preparation complete! Ready for containerd + kubeadm installation."

9.2. Run script on all nodes

# Từ workstation, distribute và chạy:
for host in master{1..3} worker{1..3} storage{1..3}; do
  echo "=== Preparing $host ==="
  scp prepare-k8s-node.sh root@${host}:/tmp/
  ssh root@${host} "bash /tmp/prepare-k8s-node.sh"
  echo "=== $host DONE ==="
  echo ""
done

PART 10: VERIFICATION CHECKLIST

10.1. Checklist for each node

#!/bin/bash
# verify-node.sh - Chạy trên mỗi node để verify
echo "=== Node: $(hostname) ==="

1. OS version

echo -n "OS: "; cat /etc/os-release | grep PRETTY_NAME | cut -d'"' -f2

2. Kernel version

echo -n "Kernel: "; uname -r

3. cgroup v2

echo -n "cgroup: "; stat -fc %T /sys/fs/cgroup/

4. Swap

SWAP=$(free -m | grep Swap | awk '{print $2}') if [ "$SWAP" -eq 0 ]; then echo "Swap: ✅ Disabled" else echo "Swap: ❌ Still enabled ($SWAP MB)" fi

5. ip_forward

FWD=$(sysctl -n net.ipv4.ip_forward) echo "ip_forward: $([ "$FWD" == "1" ] && echo "✅" || echo "❌") ($FWD)"

6. br_netfilter

BNF=$(sysctl -n net.bridge.bridge-nf-call-iptables 2>/dev/null) echo "bridge-nf-call: $([ "$BNF" == "1" ] && echo "✅" || echo "❌") ($BNF)"

7. NTP sync

SYNC=$(chronyc tracking 2>/dev/null | grep "Leap status" | awk '{print $4}') echo "NTP: $([ "$SYNC" == "Normal" ] && echo "✅" || echo "⚠️") ($SYNC)"

8. Connectivity

for target in master{1..3} worker{1..3}; do if ping -c 1 -W 1 $target &>/dev/null; then echo "Ping $target: ✅" else echo "Ping $target: ❌" fi done


💡 KEY TAKEAWAYS

  1. Kernel modules overlay and br_netfilter are required for K8s pod networking
  2. ip_forward = 1 allows pod-to-pod traffic to pass through node
  3. Swap must turn off permanently — kubelet will refuse to start if swap on
  4. NTP synchronous is critical for etcd, TLS certificates, and Ceph
  5. SSH hardening: key-based auth only, disable password login
  6. Automation script helps prepare all nodes consistently

🎯 EXERCISES

Exercise 1: Prepare all nodes

  • Run prepare-k8s-node.sh on all 7 VMs (or servers)
  • Run verify-node.sh on each node, ensuring all checks PASS
  • Screenshot of verify-node.sh results of 1 master and 1 worker

Exercise 2: Benchmark NTP

  • Check clock offset between all nodes: chronyc sources -v
  • Ensure offset < 1ms giữa tất cả nodes
  • Try interrupting NTP, change the time, and see how long it takes for chrony to re-synchronize

Exercise 3: RHEL 9 variant__HTMLTAG_344___
  • Rewrite prepare-k8s-node.sh for RHEL 9 / Rocky Linux 9
  • Replace apt → dnf, ufw → firewalld
  • Test on 1 VM to verify

📚 NEXT POST

In Lesson 4: Load Balancer for Kubernetes API Server, we will install keepalived + HAProxy to create Virtual IP for K8s API server, ensuring HA for control plane access.