Yêu cầu đầu vào:
Kiến thức cơ bản về Linux (systemd, networking, filesystem)
Hiểu biết về Docker và containerization
Kiến thức về networking cơ bản (TCP/IP, DNS, load balancing)
Hệ thống đã sử dụng cgroup v2 (yêu cầu bắt buộc từ K8s 1.36+)
containerd 2.0+ (yêu cầu bắt buộc từ K8s 1.36+)
🎯 MỤC TIÊU KHÓA HỌC
Sau khi hoàn thành khóa học, học viên sẽ:
Hiểu kiến trúc và các thành phần của Kubernetes 1.32+
Triển khai và quản lý Kubernetes cluster với containerd 2.0 và cgroup v2
Deploy và quản lý ứng dụng với các tính năng mới: Sidecar containers, In-Place Pod Resizing, Dynamic Resource Allocation
Cấu hình networking với Gateway API và Cilium (eBPF)
Thực hiện observability với OpenTelemetry, Grafana Alloy, Loki, Tempo
Áp dụng security hiện đại: ValidatingAdmissionPolicy, Pod Security Standards, Supply Chain Security
Vận hành AI/ML workloads trên Kubernetes với Dynamic Resource Allocation (DRA)
📚 NỘI DUNG KHÓA HỌC
MODULE 1: GIỚI THIỆU VÀ CƠ BẢN
Chương 1.1: Container Orchestration và Kubernetes
Container orchestration là gì và tại sao cần?
Lịch sử: Google Borg → Kubernetes (2014) → CNCF
Kubernetes 2026: Universal Control Plane — không chỉ containers, còn VMs, serverless, edge, AI pipelines
So sánh: Kubernetes vs K3s vs k0s vs Nomad (Docker Swarm không còn phù hợp production 2026)
Kubernetes ecosystem 2026: CNCF landscape, các dự án graduated
Chương 1.2: Kiến trúc Kubernetes
Control Plane components:
kube-apiserver
etcd
kube-scheduler
kube-controller-manager
cloud-controller-manager
Node components:
kubelet
kube-proxy (chú ý: IPVS mode deprecated K8s 1.35, dùng nftables)
Container runtime: containerd 2.0 (mặc định), CRI-O
Tại sao không dùng Docker làm container runtime? (dockershim đã bị xóa từ K8s 1.24)
Add-ons và plugins
Chương 1.3: Cài đặt môi trường (2026)
Yêu cầu hệ thống 2026: cgroup v2 (bắt buộc), containerd 2.0+
Các phương pháp cài đặt Kubernetes local:
Minikube (hỗ trợ containerd 2.0)
kind (Kubernetes in Docker)
k3d (K3s in Docker — nhẹ hơn, khởi động nhanh)
Cài đặt kubectl và cấu hình kubeconfig
Công cụ CLI không thể thiếu: k9s (terminal UI), kubectx/kubens, stern
Dashboard: Headlamp (thay thế chính thức cho Kubernetes Dashboard đã archived T1/2026)
IDE: Lens (personal tier miễn phí), FreeLens (open-source fork)
Bài thực hành 1:
Kiểm tra cgroup v2 và cài đặt containerd 2.0
Khởi động cluster với kind hoặc k3d
Cài đặt và cấu hình k9s, Headlamp
Chạy các lệnh kubectl cơ bản
MODULE 2: KUBERNETES OBJECTS CƠ BẢN
Chương 2.1: Pods
Pod là gì và tại sao cần Pod?
Pod lifecycle
Multi-container Pods
Init containers
Sidecar containers (GA từ K8s 1.33): init container với
restartPolicy: Always— giải quyết vấn đề lifecycle của sidecar proxy (Envoy, OTel collector, log agent)Ephemeral containers (cho debugging)
Pod templates và Static Pods
Chương 2.2: ReplicaSets và Deployments
ReplicaSet: quản lý số lượng Pod replicas
Deployment: declarative updates
Rolling updates và rollbacks
Deployment strategies: Recreate, RollingUpdate, Blue/Green, Canary
Scaling applications
Chương 2.3: Services và EndpointSlices
Service discovery trong Kubernetes
Service types: ClusterIP, NodePort, LoadBalancer, ExternalName
EndpointSlices (chuẩn mới, Endpoints API deprecated K8s 1.33)
Headless Services
Gateway API vs Ingress (xem Module 4)
Chương 2.4: Namespaces
Tổ chức resources với Namespaces
Resource quotas và Limit ranges
Network policies với Namespaces
Best practices multi-tenancy
Bài thực hành 2:
Deploy ứng dụng web với Deployment và Sidecar container (log agent)
Tạo Deployment với multiple replicas, thực hiện rolling update
Expose service với các types khác nhau
Debug với ephemeral containers (
kubectl debug)
MODULE 3: CONFIGURATION VÀ STORAGE
Chương 3.1: ConfigMaps và Secrets
Quản lý configuration với ConfigMaps
Secrets: quản lý sensitive data
Các loại Secrets; Mounting ConfigMaps và Secrets
Immutable ConfigMaps và Secrets
Secrets encryption at rest
External Secrets Operator: đồng bộ secrets từ AWS Secrets Manager, GCP Secret Manager, HashiCorp Vault
Chương 3.2: Persistent Storage
Volumes trong Kubernetes: emptyDir, hostPath
PersistentVolumes (PV), PersistentVolumeClaims (PVC), StorageClasses
Dynamic provisioning
CSI (Container Storage Interface) drivers — bắt buộc dùng CSI thay in-tree plugins (CephFS in-tree đã bị xóa K8s 1.31)
Volume snapshots và restore
VolumeAttributesClass (tính năng mới K8s 1.29+): thay đổi IOPS/throughput mà không cần xóa PVC
Chương 3.3: StatefulSets
StatefulSets vs Deployments
Stable network identities, ordered deployment, persistent storage
Headless services
Use cases: databases, distributed systems (Kafka, Zookeeper, etcd)
Bài thực hành 3:
Deploy ứng dụng với ConfigMaps và Secrets, tích hợp External Secrets Operator
Cài đặt CSI driver (ví dụ: Longhorn hoặc OpenEBS)
Deploy PostgreSQL với StatefulSet và PVC
Tạo volume snapshot và restore
MODULE 4: NETWORKING NÂNG CAO
Chương 4.1: Kubernetes Networking Model
Container-to-Container, Pod-to-Pod, Pod-to-Service, External-to-Service networking
CNI (Container Network Interface) và các lựa chọn:
Cilium (khuyến nghị 2026): eBPF-based, L7 load balancing, built-in observability với Hubble, hỗ trợ Gateway API native, không cần sidecar proxy cho service mesh
Calico: mature, eBPF dataplane, phù hợp khi cần tương thích rộng
Flannel: đơn giản nhưng thiếu Network Policy và observability — tránh dùng production
kube-proxy modes: iptables (legacy), nftables (khuyến nghị — IPVS deprecated K8s 1.35)
Chương 4.2: Gateway API (Chuẩn mới — thay thế Ingress)
Gateway API v1.4 GA (10/2025) — chuẩn mới thay thế Ingress controller truyền thống
Tại sao Gateway API tốt hơn Ingress? (role-oriented, expressive, portable)
Các resource chính: GatewayClass, Gateway, HTTPRoute, GRPCRoute, TCPRoute
BackendTLSPolicy: TLS between gateway và backend (v1.4)
Traffic splitting, header matching, URL rewriting
Implementations: Cilium Gateway API, Envoy Gateway, nginx-gateway-fabric, Istio
Ingress truyền thống: vẫn được hỗ trợ nhưng Ingress-NGINX đang vào maintenance mode (3/2026)
Chương 4.3: Network Policies
Network isolation, Pod selectors, Ingress/Egress rules
Cilium Network Policies: L7 policies (HTTP method, path, header-based)
Best practices: default-deny, least privilege
Bài thực hành 4:
Cài đặt Cilium làm CNI với Hubble UI
Cấu hình Gateway API (HTTPRoute) cho multiple services
Setup TLS với cert-manager và Gateway API
Implement L7 Network Policies với Cilium
Quan sát network traffic qua Hubble
MODULE 5: WORKLOAD MANAGEMENT
Chương 5.1: Jobs và CronJobs
Batch processing với Jobs: single, parallel, indexed, work queue
Job backoff, retries, và pod failure policies
CronJobs: scheduled tasks với timezone support (GA K8s 1.27)
JobSet (CNCF project): quản lý nhóm Jobs phụ thuộc nhau — lý tưởng cho AI/ML training pipeline
Chương 5.2: DaemonSets
DaemonSet use cases: logging agent, monitoring, network plugin
Node selection, updating DaemonSets
Ví dụ thực tế: deploy Grafana Alloy collector trên mọi node
Chương 5.3: Autoscaling
HorizontalPodAutoscaler (HPA): CPU/memory và custom metrics
VerticalPodAutoscaler (VPA): tự động điều chỉnh resource requests
In-Place Pod Resource Updates (K8s 1.35): thay đổi CPU/memory mà không cần restart Pod
KEDA (Kubernetes Event-Driven Autoscaling): scale to zero, scale dựa trên Kafka, RabbitMQ, HTTP requests, Cron...
Cluster Autoscaler: thêm/bớt nodes tự động
Karpenter (AWS/Azure): thay thế Cluster Autoscaler hiện đại hơn
Chương 5.4: Dynamic Resource Allocation (DRA) — Tính năng mới GA K8s 1.34
DRA là gì? Tại sao cần thay thế extended resources cũ?
ResourceClaim và ResourceClass
Sử dụng DRA để allocate GPU, FPGA, NIC cho AI/ML workloads
NVIDIA GPU Operator với DRA
Bài thực hành 5:
Tạo Indexed Job để process dataset song song
Cấu hình HPA với custom metrics từ Prometheus
Demo In-Place Pod resizing: thay đổi CPU limit không cần restart
Cài đặt KEDA và scale ứng dụng theo HTTP requests
MODULE 6: SECURITY
Chương 6.1: Authentication và Authorization
User authentication, ServiceAccounts
RBAC: Roles, ClusterRoles, RoleBindings, ClusterRoleBindings
Admission Controllers
Pod Security Standards (PSS) + Pod Security Admission (PSA) — thay thế PodSecurityPolicy (đã xóa từ K8s 1.25):
Privileged: không hạn chế
Baseline: ngăn leo thang đặc quyền (khuyến nghị default)
Restricted: hardened, chạy non-root
Chương 6.2: ValidatingAdmissionPolicy (GA K8s 1.30)
Tại sao cần ValidatingAdmissionPolicy? So sánh với OPA/Gatekeeper webhook
CEL (Common Expression Language) expressions
Viết policy không cần deploy webhook server
Khi nào vẫn cần OPA/Gatekeeper? (mutating, phức tạp hơn)
Chương 6.3: Security Best Practices 2026
SecurityContext: non-root user, read-only filesystem, drop capabilities
Secrets encryption at rest
Network Policies for isolation (xem Module 4)
Supply Chain Security: ký và verify container images với Cosign/Sigstore
SBOM (Software Bill of Materials)
Image pull policies và registry security
Chương 6.4: Security Tools
kube-bench: CIS Benchmark compliance
Trivy: vulnerability scanning cho images và Kubernetes manifests
Falco: runtime threat detection (process spawning bất thường, file access...)
OPA/Gatekeeper: advanced policy enforcement khi cần mutating
Bài thực hành 6:
Tạo ServiceAccounts và RBAC với least privilege
Viết ValidatingAdmissionPolicy bằng CEL (ví dụ: chặn images không có tag)
Ký container image với Cosign và verify khi deploy
Scan cluster với kube-bench và xử lý findings
Cấu hình Pod Security Admission ở chế độ Restricted
MODULE 7: OBSERVABILITY VÀ LOGGING
Chương 7.1: Observability Stack 2026 — PLG + OpenTelemetry
3 pillars của observability: Metrics, Logs, Traces
OpenTelemetry (OTel) là chuẩn duy nhất: auto-instrumentation, vendor-agnostic
Grafana Alloy: unified collector thay thế Promtail + OTel Collector + Prometheus remote-write
Stack khuyến nghị 2026:
Metrics: Prometheus + kube-state-metrics + node exporter
Logs: Loki (thay thế Elasticsearch cho logs — nhẹ hơn, rẻ hơn)
Traces: Tempo
Visualization: Grafana
Collector: Grafana Alloy
EFK Stack (Elasticsearch + Fluentd + Kibana): vẫn viable cho full-text search nhưng nặng hơn PLG
Chương 7.2: Prometheus và Grafana
Prometheus Operator và kube-prometheus-stack
ServiceMonitor và PodMonitor
Grafana dashboards: Kubernetes cluster, nodes, workloads
AlertManager: rules, routing, receivers (Slack, PagerDuty, email)
Recording rules và best practices
Chương 7.3: Loki, Tempo và Distributed Tracing
Loki: log aggregation với label-based querying (LogQL)
Grafana Alloy thu thập logs từ containers
Tempo: distributed tracing
Kết hợp Logs + Traces + Metrics trong Grafana (correlated observability)
Chương 7.4: Debugging và Troubleshooting
kubectl commands:
kubectl debug,kubectl events,kubectl topEphemeral containers để debug running pods
Node issues:
kubectl describe node, node conditionsNetwork debugging với Cilium Hubble
Performance troubleshooting, common issues và solutions
Bài thực hành 7:
Deploy kube-prometheus-stack (Prometheus + Grafana + AlertManager)
Deploy Loki + Grafana Alloy để collect logs
Deploy Tempo và cấu hình OpenTelemetry auto-instrumentation cho ứng dụng
Tạo Grafana dashboard với correlated metrics, logs, traces
Tạo alerting rule và test notification
MODULE 8: ADVANCED TOPICS
Chương 8.1: Helm 4
Helm architecture và Helm 4 vs Helm 3 (release T11/2025 — 10th anniversary)
Helm 4 features mới: WebAssembly (WASM) plugins, server-side apply, 60% performance improvement, OCI enhancements
Charts structure, installing và creating custom charts
Chart repositories và OCI registry
Helm hooks, tests, Helmfile
Helm 3 vẫn nhận security fixes đến T11/2026
Chương 8.2: Operators và Custom Resources
Operator pattern và use cases
Custom Resource Definitions (CRDs) và Custom Controllers
Operator SDK và Kubebuilder
Common operators: Prometheus Operator, CloudNativePG (PostgreSQL), Strimzi (Kafka)
Chương 8.3: Service Mesh 2026
Tại sao cần Service Mesh? mTLS, traffic management, observability
Cilium Service Mesh (Sidecarless): eBPF ở kernel layer, không cần sidecar proxy, 40-60% giảm network overhead
Istio: tính năng đầy đủ nhất, phù hợp enterprise (multi-cluster, granular RBAC)
Linkerd: nhẹ nhất, Rust-based micro-proxy, lý tưởng cho resource-constrained environments
Khi nào chọn gì? Comparison chart
Chương 8.4: GitOps
GitOps principles: Git là single source of truth
ArgoCD 3.x: centralized, hub-and-spoke multi-cluster, single pane of glass
Flux 2.x: decentralized, cluster tự pull từ Git/OCI, an toàn hơn cho distributed teams
Chọn ArgoCD hay Flux? Architectural tradeoffs
CI/CD pipeline với GitHub Actions + ArgoCD/Flux
Bài thực hành 8:
Tạo Helm chart với Helm 4, publish lên OCI registry
Build Operator đơn giản với Kubebuilder
So sánh Cilium Service Mesh vs Istio trên cùng cluster
Setup GitOps với ArgoCD: deploy ứng dụng từ Git
MODULE 9: CLUSTER MANAGEMENT
Chương 9.1: Production Cluster Setup
kubeadm installation với containerd 2.0 và cgroup v2
High Availability: multi-master architecture, load balancing
kube-proxy: cấu hình nftables mode (IPVS deprecated K8s 1.35)
Cluster upgrades: chiến lược upgrade an toàn từng minor version
Backup với Velero: cluster state và PV snapshots
Chương 9.2: Infrastructure Migration (quan trọng 2026)
Migrate cgroup v1 → cgroup v2: bắt buộc trước khi upgrade lên K8s 1.36
Upgrade containerd 1.x → containerd 2.0: bắt buộc từ K8s 1.36
Kiểm tra compatibility của workloads với cgroup v2
Checklist migration production cluster
Chương 9.3: Node Management
Adding/removing nodes
Node maintenance: drain, cordon, uncordon
Taints và Tolerations
Node affinity, anti-affinity, topology spread constraints
Pod priority và preemption
Chương 9.4: Resource Management
Resource requests và limits
Quality of Service (QoS) classes: Guaranteed, Burstable, BestEffort
LimitRanges và ResourceQuotas
Pod Disruption Budgets (PDB)
Cluster Autoscaler vs Karpenter
Chương 9.5: Cluster API
Cluster API là gì? Declarative cluster lifecycle management
Infrastructure providers: AWS, GCP, Azure, vSphere
Tạo và upgrade clusters bằng kubectl
Bài thực hành 9:
Setup 3-node cluster với kubeadm + containerd 2.0 + cgroup v2
Thực hiện migrate cgroup v1 → v2 trên cluster đang chạy
Perform cluster upgrade từ 1.32 → 1.33
Cài đặt Velero và thực hiện backup/restore
Practice disaster recovery scenario
MODULE 10: CLOUD PLATFORMS VÀ BEST PRACTICES
Chương 10.1: Managed Kubernetes Services 2026
Amazon EKS: Auto Mode, Pod Identity, EKS Anywhere
Google GKE: Autopilot, Workload Identity, GKE Enterprise
Azure AKS: Automatic upgrades, KEDA integration, Workload Identity
So sánh pricing, features, và tính năng riêng của mỗi platform
Chương 10.2: Cost Optimization
Right-sizing với VPA recommendations
Spot/Preemptible instances cho workloads tolerant
Karpenter: node consolidation và Spot interruption handling
Kubecost hoặc OpenCost: visibility vào chi phí per namespace/team
Chương 10.3: Best Practices Production 2026
Cluster setup: multi-AZ, control plane HA
Application: resource limits, liveness/readiness probes, PDB
Security hardening theo CIS Kubernetes Benchmark
Multi-tenancy: Namespace isolation, Hierarchical Namespaces Controller
Hybrid và multi-cloud: Cluster Federation, Multi-cluster Gateway
Edge computing với K3s (lightweigh, ARM support)
Bài thực hành 10:
Deploy production workload trên GKE Autopilot hoặc EKS Auto Mode
Cài đặt OpenCost và phân tích chi phí theo team
Implement Karpenter với Spot instances
Final project: deploy complete microservices application với Gateway API, Cilium, GitOps, Observability stack
MODULE 11: AI/ML WORKLOADS TRÊN KUBERNETES (2026)
Chương 11.1: Kubernetes cho AI/ML
Tại sao Kubernetes là platform lý tưởng cho AI/ML workloads?
GPU support: NVIDIA GPU Operator, cài đặt và cấu hình
Dynamic Resource Allocation (DRA) GA K8s 1.34: GPU sharing, FPGA allocation
Node selectors và taints/tolerations cho GPU nodes
Resource quotas cho GPU workloads
Chương 11.2: Training Jobs
JobSet: điều phối nhiều Jobs phụ thuộc nhau trong training pipeline
Kubeflow Training Operator: PyTorchJob, TFJob, MXJob
Distributed training patterns: Data parallelism, Model parallelism
Checkpoint và resume training
Chương 11.3: Model Serving và Inference
Kubernetes Inference Extension (KIE): chuẩn mới cho LLM serving
KServe: model serving framework (TensorFlow, PyTorch, ONNX...)
vLLM trên Kubernetes cho LLM inference
Autoscaling inference với KEDA (scale dựa trên queue depth, GPU utilization)
Chương 11.4: MLOps Pipelines
Kubeflow Pipelines: orchestrate ML workflows
Argo Workflows: general-purpose workflow engine
Data processing với Spark on Kubernetes
Model registry và versioning
Bài thực hành 11:
Cài đặt NVIDIA GPU Operator và verify GPU scheduling
Tạo PyTorchJob với Kubeflow Training Operator
Deploy LLM inference server với KServe hoặc vLLM
Cấu hình KEDA autoscaling cho inference workload
📖 TÀI LIỆU THAM KHẢO
Tài liệu chính thức
Kubernetes Official Documentation: https://kubernetes.io/docs/
Kubernetes Blog: https://kubernetes.io/blog/
CNCF Projects: https://www.cncf.io/projects/
Gateway API: https://gateway-api.sigs.k8s.io/
OpenTelemetry: https://opentelemetry.io/
Sách
"Kubernetes Up & Running" 3rd Ed — Kelsey Hightower (cập nhật 2023)
"The Kubernetes Book" — Nigel Poulton (cập nhật hàng năm)
"Kubernetes Patterns" 2nd Ed — Bilgin Ibryam & Roland Huß
"Production Kubernetes" — Josh Rosso et al.
"Cloud Native Observability with OpenTelemetry" — Alex Boten
Khóa học và chứng chỉ
Certified Kubernetes Administrator (CKA) — Linux Foundation
Certified Kubernetes Application Developer (CKAD) — Linux Foundation
Certified Kubernetes Security Specialist (CKS) — Linux Foundation
Labs và Practice
KillerCoda: https://killercoda.com/ (thay thế Katacoda đã đóng cửa 2023)
killer.sh: môi trường luyện thi CKA/CKAD/CKS
Kubernetes the Hard Way (Kelsey Hightower)
Play with Kubernetes: https://labs.play-with-k8s.com/
Community
Kubernetes Slack: https://slack.k8s.io/
Kubernetes GitHub: https://github.com/kubernetes/kubernetes
KubeCon + CloudNativeCon conferences
🔧 CÔNG CỤ CẦN THIẾT
Essential Tools
kubectl
kubeadm
kind / k3d / Minikube
containerd 2.0+ (thay Docker daemon làm runtime)
VS Code với Kubernetes extension + YAML extension
Recommended CLI Tools
k9s — terminal UI, cực nhanh, "Vim của Kubernetes"
kubectx / kubens — switch context/namespace
stern — multi-pod log streaming
Helm 4 — package manager
kustomize — built-in kubectl, không cần cài thêm
cilium CLI — quản lý và debug Cilium
Dashboard và IDE
Headlamp — web UI chính thức thay thế Kubernetes Dashboard (archived T1/2026), được SIG UI endorsement
K9s — terminal UI cho power users
Lens — desktop IDE, personal tier miễn phí; enterprise tier trả phí
FreeLens — open-source fork của Lens (OpenLens không còn được maintain)
Observability Stack
Prometheus + Grafana + AlertManager
Loki (logs) + Tempo (traces) + Grafana Alloy (collector)
OpenTelemetry Operator
Hubble UI (Cilium network observability)
🎊 KẾT LUẬN
Kubernetes năm 2026 đã trưởng thành vượt bậc: từ một công cụ container orchestration, K8s đang trở thành Universal Control Plane cho mọi workload — containers, VMs, AI/ML pipelines, edge devices. Những tính năng như Gateway API, Cilium eBPF, Sidecar containers GA, ValidatingAdmissionPolicy, và Dynamic Resource Allocation cho thấy hệ sinh thái đang ngày càng mạnh mẽ và production-ready hơn.
Quan trọng nhất: thực hành thường xuyên, theo dõi Kubernetes Blog và CNCF landscape để không bị lạc hậu trong một ecosystem phát triển nhanh như vậy.
Lưu ý: Khóa học này được thiết kế dựa trên Kubernetes version 1.32+ (phiên bản LTS phù hợp production đầu 2026). Phiên bản mới nhất tại thời điểm viết là Kubernetes 1.35.3 (tháng 3/2026). Kiểm tra release notes tại kubernetes.io/releases trước khi upgrade production cluster.
