Chuyển đến nội dung chính

LESSON 7: JOIN ADD CONTROL PLANE AND WORKER NODES

Join master2 and master3 to HA control plane, join worker nodes, verify etcd cluster with 3 members, check leader election and cluster readiness.

🔒 DevSecOps — Lesson 7 LESSON 7: JOIN ADDING CONTROL PLANE AND WORKER NODES

Deploy Microservices On-Premises with Kubernetes HA

Part 2: Kubernetes HA Cluster with kubeadm__HTMLTAG_62___

xdev.asia

🎯 LESSON OBJECTIVE__HTMLTAG_68___

After completing this lesson, you will:

  • ✅ Join 2 more control plane nodes to form an HA cluster of 3 masters
  • ✅ Join worker nodes into cluster__HTMLTAG_75___
  • ✅ Verify etcd cluster 3 members works correctly
  • ✅ Check leader election for scheduler and controller-manager__HTMLTAG_79___
  • ✅ Label and taint nodes according to the correct role

PART 1: JOIN CONTROL PLANE NODES

1.1. Prepare on master2 and master3

Make sure master2 and master3 are completed:

  • ✅ OS tuning (Lesson 3)
  • ✅ containerd + kubeadm install (Lesson 5)
  • ✅ Can connect to VIP 10.10.20.100:6443
# Trên master2 và master3, verify connectivity:
nc -zv 10.10.20.100 6443
# Connection to 10.10.20.100 6443 port [tcp/*] succeeded!

# Verify containerd running:
systemctl status containerd
# ● containerd.service - containerd container runtime
#    Active: active (running)

1.2. Join master2 into Control Plane

# Trên master2:
# Tạo audit log directory trước:
mkdir -p /var/log/kubernetes

Join command (thay <TOKEN>, <HASH>, <CERT_KEY> từ output Bài 6):

kubeadm join 10.10.20.100:6443
--token <TOKEN>
--discovery-token-ca-cert-hash sha256:<HASH>
--control-plane
--certificate-key <CERT_KEY>
--apiserver-advertise-address 10.10.20.12

Output:

[preflight] Running pre-flight checks

[preflight] Reading configuration from the cluster...

[download-certs] Downloading the certificates in Secret "kubeadm-certs"

[certs] Using certificateDir folder "/etc/kubernetes/pki"

[certs] Generating "apiserver" certificate and key

...

[mark-control-plane] Marking the node master2 as control-plane

This node has joined the cluster and a new control plane instance was created.

Run 'kubectl get nodes' on any control-plane node to see this node join.

⚠️ --apiserver-advertise-address: Each master uses its own IP on the cluster network.

1.3. Join master3 into Control Plane

# Trên master3:
mkdir -p /var/log/kubernetes

kubeadm join 10.10.20.100:6443
--token <TOKEN>
--discovery-token-ca-cert-hash sha256:<HASH>
--control-plane
--certificate-key <CERT_KEY>
--apiserver-advertise-address 10.10.20.13

1.4. Setup kubeconfig on master2 & master3

# Trên mỗi master node (master2 và master3):
mkdir -p $HOME/.kube
cp -i /etc/kubernetes/admin.conf $HOME/.kube/config
chown $(id -u):$(id -g) $HOME/.kube/config

Verify nodes:

kubectl get nodes

NAME STATUS ROLES AGE VERSION

master1 NotReady control-plane 10m v1.31.0

master2 NotReady control-plane 3m v1.31.0

master3 NotReady control-plane 1m v1.31.0


PART 2: HANDLING EXPIRED TOKENS

2.1. Create new token (if token expires)

# Token hết hạn sau 24 giờ. Tạo mới:
kubeadm token create --print-join-command
# Output:
# kubeadm join 10.10.20.100:6443 --token NEW_TOKEN --discovery-token-ca-cert-hash sha256:HASH

Certificate-key hết hạn sau 2 giờ. Upload lại:

kubeadm init phase upload-certs --upload-certs

Output:

[upload-certs] Using certificate key: NEW_CERT_KEY

Kết hợp lại để join control-plane:

kubeadm join 10.10.20.100:6443 \

--token NEW_TOKEN \

--discovery-token-ca-cert-hash sha256:HASH \

--control-plane \

--certificate-key NEW_CERT_KEY


PART 3: JOIN WORKER NODES

3.1. Join worker1, worker2, worker3

# Trên MỖI worker node (worker1, worker2, worker3):
kubeadm join 10.10.20.100:6443 \
  --token <TOKEN> \
  --discovery-token-ca-cert-hash sha256:<HASH>

Output:

[preflight] Running pre-flight checks

[preflight] Reading configuration from the cluster...

[kubelet-start] Starting the kubelet

[kubelet-start] Waiting for the kubelet to perform the TLS Bootstrap

This node has joined the cluster:

* Certificate signing request was sent to apiserver and a response was received.

* The Kubelet was informed of the new secure connection details.

3.2. Verify all nodes

# Trên bất kỳ master nào:
kubectl get nodes -o wide
# NAME      STATUS     ROLES           AGE   VERSION    INTERNAL-IP    OS-IMAGE
# master1   NotReady   control-plane   30m   v1.31.0    10.10.20.11    Ubuntu 24.04 LTS
# master2   NotReady   control-plane   20m   v1.31.0    10.10.20.12    Ubuntu 24.04 LTS
# master3   NotReady   control-plane   18m   v1.31.0    10.10.20.13    Ubuntu 24.04 LTS
# worker1   NotReady   <none>          5m    v1.31.0    10.10.20.21    Ubuntu 24.04 LTS
# worker2   NotReady   <none>          4m    v1.31.0    10.10.20.22    Ubuntu 24.04 LTS
# worker3   NotReady   <none>          3m    v1.31.0    10.10.20.23    Ubuntu 24.04 LTS
# ← Status = NotReady vì chưa cài CNI (Bài 8: Cilium)

PART 4: VERIFY ETCD CLUSTER

4.1. etcd member list

# Kiểm tra etcd cluster membership:
kubectl -n kube-system exec etcd-master1 -- etcdctl \
  --endpoints=https://127.0.0.1:2379 \
  --cacert=/etc/kubernetes/pki/etcd/ca.crt \
  --cert=/etc/kubernetes/pki/etcd/server.crt \
  --key=/etc/kubernetes/pki/etcd/server.key \
  member list -w table

Output:

+------------------+---------+---------+----------------------------+----------------------------+

| ID | STATUS | NAME | PEER ADDRS | CLIENT ADDRS |

+------------------+---------+---------+----------------------------+----------------------------+

| 1a2b3c4d5e6f7890 | started | master1 | https://10.10.20.11:2380 | https://10.10.20.11:2379 |

| 2b3c4d5e6f789012 | started | master2 | https://10.10.20.12:2380 | https://10.10.20.12:2379 |

| 3c4d5e6f78901234 | started | master3 | https://10.10.20.13:2380 | https://10.10.20.13:2379 |

+------------------+---------+---------+----------------------------+----------------------------+

4.2. etcd endpoint health

# Health check tất cả endpoints:
kubectl -n kube-system exec etcd-master1 -- etcdctl \
  --endpoints=https://10.10.20.11:2379,https://10.10.20.12:2379,https://10.10.20.13:2379 \
  --cacert=/etc/kubernetes/pki/etcd/ca.crt \
  --cert=/etc/kubernetes/pki/etcd/server.crt \
  --key=/etc/kubernetes/pki/etcd/server.key \
  endpoint health -w table

Output:

+----------------------------+--------+-------------+-------+

| ENDPOINT | HEALTH | TOOK | ERROR |

+----------------------------+--------+-------------+-------+

| https://10.10.20.11:2379 | true | 10.123456ms | |

| https://10.10.20.12:2379 | true | 12.345678ms | |

| https://10.10.20.13:2379 | true | 11.234567ms | |

+----------------------------+--------+-------------+-------+

4.3. etcd endpoint status (leader check)

# Check leader:
kubectl -n kube-system exec etcd-master1 -- etcdctl \
  --endpoints=https://10.10.20.11:2379,https://10.10.20.12:2379,https://10.10.20.13:2379 \
  --cacert=/etc/kubernetes/pki/etcd/ca.crt \
  --cert=/etc/kubernetes/pki/etcd/server.crt \
  --key=/etc/kubernetes/pki/etcd/server.key \
  endpoint status -w table

Output:

+----------------------------+------------------+---------+---------+-----------+...+--------+

| ENDPOINT | ID | VERSION | DB SIZE | IS LEADER |...| ERRORS |

+----------------------------+------------------+---------+---------+-----------+...+--------+

| https://10.10.20.11:2379 | 1a2b3c4d5e6f7890 | 3.5.15 | 3.3 MB | true |...| |

| https://10.10.20.12:2379 | 2b3c4d5e6f789012 | 3.5.15 | 3.3 MB | false |...| |

| https://10.10.20.13:2379 | 3c4d5e6f78901234 | 3.5.15 | 3.3 MB | false |...| |

+----------------------------+------------------+---------+---------+-----------+...+--------+


PART 5: LEADER ELECTION VERIFICATION

5.1. Check Controller Manager leader

# Controller Manager sử dụng Lease để bầu leader:
kubectl -n kube-system get lease kube-controller-manager -o yaml
# holderIdentity: master1_xxxxx  ← master1 là leader hiện tại
# leaseDurationSeconds: 15
# renewTime: "2025-04-02T07:00:30Z"

Scheduler leader:

kubectl -n kube-system get lease kube-scheduler -o yaml

holderIdentity: master1_xxxxx ← master1 là leader hiện tại

5.2. Test HA Failover

# Test: Shutdown master1, verify cluster vẫn hoạt động
# ⚠️ Chỉ test trong lab!

Trên master1:

systemctl stop kubelet

Trên master2, kiểm tra:

kubectl get nodes

master1 sẽ chuyển sang NotReady sau ~40 giây

Controller Manager và Scheduler leader sẽ tự chuyển sang master2 hoặc master3

Verify scheduler leader changed:

kubectl -n kube-system get lease kube-scheduler -o jsonpath='{.spec.holderIdentity}'

master2_xxxxx hoặc master3_xxxxx

Khôi phục master1:

systemctl start kubelet

master1 sẽ trở lại Ready sau vài giây


PART 6: LABEL AND TAINT NODES

6.1. Label worker nodes

# Gán role cho worker nodes (mặc định workers không có role label):
kubectl label node worker1 node-role.kubernetes.io/worker=""
kubectl label node worker2 node-role.kubernetes.io/worker=""
kubectl label node worker3 node-role.kubernetes.io/worker=""

Gán labels cho topology:

kubectl label node worker1 topology.kubernetes.io/zone=rack-a kubectl label node worker2 topology.kubernetes.io/zone=rack-b kubectl label node worker3 topology.kubernetes.io/zone=rack-a

Label cho Ceph storage nodes (nếu dùng):

kubectl label node worker1 storage-node=true kubectl label node worker2 storage-node=true kubectl label node worker3 storage-node=true

Verify labels:

kubectl get nodes --show-labels

6.2. Check taints

# Control plane nodes mặc định có taint:
kubectl describe node master1 | grep -i taint
# Taints: node-role.kubernetes.io/control-plane:NoSchedule

Worker nodes KHÔNG có taint:

kubectl describe node worker1 | grep -i taint

Taints: <none>

⚠️ KHÔNG remove taint trên control-plane trong production

Control plane nodes chỉ chạy system components


PART 7: VERIFY HAPROXY BACKENDS

# Kiểm tra HAProxy stats — giờ cả 3 masters đều UP:
curl -s http://lb1:9000/stats\;csv | grep apiserver
# k8s-api,master1,... UP ...
# k8s-api,master2,... UP ...
# k8s-api,master3,... UP ...

# Verify load balancing:
for i in {1..10}; do
  curl -sk https://10.10.20.100:6443/healthz
  echo
done
# Mỗi request được HAProxy phân phối tới 1 trong 3 masters

💡 KEY TAKEAWAYS

  1. Join control-planeneed to add --control-plane --certificate-key
  2. Token expires in 24 hours, certificate-key expires in 2 hours — create new if necessary
  3. etcd 3 members allows the cluster to withstand 1 member down (quorum = 2)
  4. Leader electionautomatic: scheduler and controller-manager failover when leader is down
  5. Label and taint nodes have the right role to help with accurate scheduling
  6. NotReady status is normal — need to install CNI (Lesson 8) to convert to Ready

🎯 EXERCISES__HTMLTAG_179___

Exercise 1: Join full cluster__HTMLTAG_181___
  • Join master2 and master3 to the control plane__HTMLTAG_184___
  • Join worker1, worker2, worker3
  • Verify kubectl get nodes shows 6 nodes

Exercise 2: etcd Health Check

  • Run etcdctl member list, endpoint health, endpoint status
  • Determine who is etcd leader__HTMLTAG_196___

Exercise 3: HA Failover Test

  • Stop kubelet on master1
  • Verify cluster is still active via VIP
  • Verify leader election to another master__HTMLTAG_206___
  • Start kubelet again, verify master1 returns__HTMLTAG_208___

📚 NEXT POST

In Lesson 8: Installing Cilium CNI — eBPF Networking, we will install Cilium as a Container Network Interface, resolve NotReady status and enable NetworkPolicy.