Chuyển đến nội dung chính

Kubernetes: From Basics to Advanced

Comprehensive basic to advanced Kubernetes course for 2026, helping you master container orchestration, deploy production-ready applications, and prepare for CKA/CKAD certification. Updated to Kubernetes 1.32+ with Gateway API, Cilium, OpenTelemetry, Helm 4, Sidecar containers GA, ValidatingAdmissionPolicy, and AI/ML workloads.

Input requirements:

  • Basic knowledge of Linux (systemd, networking, filesystem)

  • Understanding about Docker and containerization

  • Basic networking knowledge (TCP/IP, DNS, load balancing)

  • The system has used cgroup v2 (mandatory requirement from K8s 1.36+)

  • containerd 2.0+ (required from K8s 1.36+)


🎯 COURSE OBJECTIVES

After completing the course, students will:

  • Understand the architecture and components of Kubernetes 1.32+_

  • Deploy and manage Kubernetes clusters with containerd 2.0 and cgroup v2_

  • Deploy and manage applications with new features: Sidecar containers, In-Place Pod Resizing, Dynamic Resource Allocation_

  • Configure networking with API Gateway and Cilium (eBPF)

  • Implement observability with OpenTelemetry, Grafana Alloy, Loki, Tempo_

  • Apply modern security: ValidatingAdmissionPolicy, Pod Security Standards, Supply Chain Security

  • Operate AI/ML workloads on Kubernetes with Dynamic Resource Allocation (DRA)


📚 KEY CONTENT LEARNING

MODULE 1: INTRODUCTION AND BASICS

Chapter 1.1: Container Orchestration and Kubernetes

  • What is container orchestration and why is it needed?

  • History: Google Borg → Kubernetes (2014) → CNCF

  • Kubernetes 2026: Universal Control Plane — not just containers, also VMs, serverless, edges, AI pipelines_

  • Comparison: Kubernetes vs K3s vs k0s vs Nomad (Docker Swarm is no longer suitable for production 2026)

  • Kubernetes ecosystem 2026: CNCF landscape, graduated projects

Chapter 1.2: Kubernetes Architecture

  • Control Plane components:

    • kube-apiserver

    • ___HTMLTA G_103___

      etcd

    • kube-scheduler

      ___HT MLTAG_110___
    • kube-controller-manager

    • ___HTMLTAG_11 6___cloud-controller-manager_

  • Node components:

    • kubelet_

    • kube-proxy (note: IPVS mode deprecated K8s 1.35, use nftables)

    • Container runtime: containerd 2.0 (default), CRI-O

  • Why not use Docker as a runtime container? (dockershim removed from K8s 1.24)_

  • Add-ons and plugins_

Chapter 1.3: Environment Settings (2026)

  • System requirements 2026: cgroup v2 (required), containerd 2.0+

  • Installation methods set Kubernetes local:

    • Minikube (containerd 2.0 support)

    • kind (Kubernetes in Docker)

    • k3d (K3s in Docker — lighter, starter quick)

  • Install kubectl and configure kubeconfig

  • Public Essential CLI tools: k9s (terminal UI), kubectx/kubens, stern_

  • Dashboard: Headlamp (official replacement for archived Kubernetes Dashboard January 2026)

  • IDE: Lens (free personal tier), FreeLens (open-source fork)

Practice 1:

  • Check cgroup v2 and install containerd 2.0

  • Start cluster with kind or k3d_

  • Install and configure k9s, Headlamp_

  • Run basic kubectl commands_


_MODULE 2: BASIC KUBERNETES OBJECTS_

Chapter 2.1: Pods_

  • _What is a Pod and why is it needed? Pod?

  • Pod lifecycle

  • Multi-container Pods

  • Init containers

  • Sidecar containers (GA from K8s 1.33): init container with restartPolicy: Always — solves sidecar proxy lifecycle problem (Envoy, OTel collector, log agent)

  • Ephemeral containers (for debugging)

  • Pod templates and Static Pods

Chapter 2.2: ReplicaSets and Deployments

  • ReplicaSet: manage the number of Pod replicas

  • Deployment: declarative updates

  • Rolling updates and rollbacks

  • Deployment strategies: Recreate, RollingUpdate, Blue/Green, Canary_

  • Scaling applications_

Chapter 2.3: Services and EndpointSlices

  • Service discovery in Kubernetes_

  • Service types: ClusterIP, NodePort, LoadBalancer, ExternalName

  • _EndpointSlices (new standard, Endpoints API deprecated K8s 1.33)

  • Headless Services_

  • _Gateway API vs Ingress (see Module 4)_

Chapter 2.4: Namespaces

  • Organize resources with Namespaces

  • Resource quotas and Limit ranges

  • Network policies with Namespaces

  • Best practices multi-tenancy

Practice 2:

  • Deploy web application with Deployment and Sidecar container (log agent)

  • Create Deployment with multiple replicas, perform rolling update_

  • Expose service with different types_

  • Debug with ephemeral containers (kubectl debug)


MODULE 3: CONFIGURATION AND STORAGE

Chapter 3.1: ConfigMaps and Secrets

  • Manage configuration with ConfigMaps

  • Secrets: manage sensitive data

  • Types of Secrets; Mounting ConfigMaps and Secrets_

  • Immutable ConfigMaps and Secrets_

  • Secrets encryption at rest

  • External Secrets Operator: synchronize secrets from AWS Secrets Manager, GCP Secret Manager, HashiCorp Vault

Chapter 3.2: Persistent Storage

  • Volumes in Kubernetes: emptyDir, hostPath_

  • PersistentVolumes (PV), PersistentVolumeClaims (PVC), StorageClasses

  • Dynamic provisioning

  • CSI (Container Storage Interface) drivers — start forced to use CSI instead of in-tree plugins (CephFS in-tree has been removed in K8s 1.31)

  • Volume snapshots and restore

  • _VolumeAttributesClass (new feature K8s 1.29+): change IOPS/throughput without removing PVC

Chapter 3.3: StatefulSets

  • StatefulSets vs Deployments_

  • Stable network identities, ordered deployment, persistent storage_

  • _Headless services

  • Use cases: databases, distributed systems (Kafka, Zookeeper, etcd)

Practice 3:

  • Deploy application with ConfigMaps and Secrets, integrate External Secrets Operator_

  • Install CSI driver (e.g. Longhorn or OpenEBS)_

  • Deploy PostgreSQL with StatefulSet and PVC_

  • Create volume snapshot and restore


MODULE 4: ADVANCED NETWORKING

Chapter 4.1: Kubernetes Networking Model

  • Container-to-Container, Pod-to-Pod, Pod-to-Service, External-to-Service networking_

  • CNI (Container Network Interface) and options:

    • _Cilium (recommended 2026): eBPF-based, L7 load balancing, built-in observability with Hubble, native Gateway API support, no need for sidecar proxy for service mesh

    • Calico_: mature, eBPF dataplane, suitable when compatibility is needed wide

    • Flannel: simple but lacks Network Policy and observability — avoid using production

  • kube-proxy modes: iptables (legacy), nftables (recommended — IPVS deprecated K8s 1.35)

Chapter 4.2: Gateway API (New standard — replacement Ingress)

  • Gateway API v1.4 GA (October 2025) — new standard replacing Ingress controller transmission system

  • Why API Gateway is better than Ingress? (role-oriented, expressive, portable)

  • Main resources: GatewayClass, Gateway, HTTPRoute, GRPCRoute, TCPRoute

  • BackendTLSPolicy: TLS between gateway and backend (v1.4)

  • Traffic splitting, header matching, URL rewriting_

  • Implementations: Cilium Gateway API, Envoy Gateway, nginx-gateway-fabric, Istio

  • Traditional Ingress: still supported but Ingress-NGINX is in maintenance mode (March 2026)

Chapter 4.3: Network Policies_

  • Network isolation, Pod selectors, Ingress/Egress rules

  • Cilium Network Policies: L7 policies (HTTP method, path, header-based)_

  • Best practices: default-deny, least privilege

_Practice 4:

  • Install Cilium as CNI with Hubble UI

  • Configure Gateway API (HTTPRoute) for multiple services_

  • Setup TLS with cert-manager and Gateway API_

  • Implement L7 Network Policies with Cilium_

  • Observe network traffic passing through Hubble


MODULE 5: WORKLOAD MANAGEMENT_

Chapter 5.1: Jobs and CronJobs_

  • Batch processing with Jobs: single, parallel, indexed, work queue_

  • Job backoff, retries, and pod failures policies

  • CronJobs: scheduled tasks with timezone support (GA K8s 1.27)

  • JobSet_ (CNCF project): manage groups of dependent Jobs — ideal for AI/ML training pipeline

Chapter 5.2: DaemonSets

  • DaemonSet use cases: logging agent, monitoring, network plugin

  • Node selection, updating DaemonSets

  • Practical example: deploy Grafana Alloy collector on all node_

Chapter 5.3: Autoscaling

  • _HorizontalPodAutoscaler (HPA): CPU/memory and custom metrics

  • VerticalPodAutoscaler (VPA): automatically adjust resources requests

  • In-Place Pod Resource Updates (K8s 1.35): change CPU/memory without restarting Pod

  • KEDA (Kubernetes Event-Driven Autoscaling): scale to zero, scale based on Kafka, RabbitMQ, HTTP requests, Cron...

  • _Cluster Autoscaler: add/remove nodes automatically_

  • _Karpenter (AWS/Azure): more modern Cluster Autoscaler replacement_

Chapter 5.4: Dynamic Resource Allocation (DRA) — New features GA K8s 1.34

  • What is DRA? Why do we need to replace old extended resources?_

  • _ResourceClaim and ResourceClass_

  • Use DRA to allocate GPU, FPGA, NIC for AI/ML workloads_

  • NVIDIA GPU Operator with DRA_

Practice 5:

  • Create Indexed Job to process dataset in parallel_

  • Configure HPA with custom metrics from Prometheus_

  • Demo In-Place Pod resizing: change CPU limit without restart_

  • _Install KEDA and scale the application via HTTP requests


MODULE 6: SECURITY

Chapter 6.1: Authentication and Authorization

  • User authentication, ServiceAccounts

  • RBAC: Roles, ClusterRoles, RoleBindings, ClusterRoleBindings_

  • Admission Controllers

  • _Pod Security Standards (PSS) + Pod Security Admission (PSA) — replaces PodSecurityPolicy (removed from K8s 1.25):

    • Privileged: unlimited mechanism

    • Baseline: prevent escalation of privilege (recommended default)

    • Restricted: hardened, run non-root

Chapter 6.2: ValidatingAdmissionPolicy (GA K8s 1.30)

  • Why do you need ValidatingAdmissionPolicy? Compare with OPA/Gatekeeper webhook

  • CEL (Common Expression Language) expressions

  • Write policies without deploying webhooks server

  • When is OPA/Gatekeeper still needed? (mutating, more complicated)

Chapter 6.3: Security Best Practices 2026_

  • _SecurityContext: non-root user, read-only filesystem, drop capabilities

  • Secrets encryption at rest

  • Network Policies for isolation (see Module 4)

  • Supply Chain Security: sign and verify container images with Cosign/Sigstore

  • SBOM (Software Bill of Materials)

  • Image pull policies and registry security

Chapter 6.4: Security Tools

  • kube-bench: CIS Benchmark compliance

  • Trivy: vulnerability scanning for images and Kubernetes manifests

  • Falco: runtime threat detection (abnormal process spawning, file access...)

  • OPA/Gatekeeper: advanced policy enforcement when necessary mutating_

Practice 6:

  • Create ServiceAccounts and RBAC with least privilege

  • Write ValidatingAdmissionPolicy with CEL (eg: block images without tags)

  • Sign container images with Cosign and verify when deploying

  • Scan cluster with kube-bench and process findings_

  • Configure Pod Security Admission in Restricted mode_


_MODULE 7: OBSERVABILITY AND LOGGING

Chapter 7.1: Observability Stack 2026 — PLG + OpenTelemetry

  • 3 pillars of observability: Metrics, Logs, Traces

  • OpenTelemetry (OTel) is the only standard: auto-instrumentation, vendor-agnostic

  • Grafana Alloy_: unified collector replacing Promtail + OTel Collector + Prometheus remote-write_

  • Stack recommended 2026:

    • _Metrics: Prometheus_ + kube-state-metrics + node exporter

    • Logs: Loki (alternative to Elasticsearch for logs — lighter, cheaper more)_

    • Traces: Tempo_

    • Visualization: Grafana

    • Collector: Grafana Alloy

  • EFK Stack (Elasticsearch + Fluentd + Kibana): still viable for full-text search but heavier than PLG

Chapter 7.2: Prometheus and Grafana

  • Prometheus Operator and kube-prometheus-stack

  • ServiceMonitor and PodMonitor_

  • Grafana dashboards: Kubernetes cluster, nodes, workloads

  • AlertManager: rules, routing, receivers (Slack, PagerDuty, email)

  • Recording rules and best practices

Chapter 7.3: Loki, Tempo and Distributed Tracing_

  • Loki: log aggregation with label-based querying (LogQL)

  • Grafana Alloy collects logs from containers_

  • Tempo: distributed tracing_

  • Combining Logs + Traces + Metrics in Grafana (correlated observability)

Chapter 7.4: Debugging and Troubleshooting

  • kubectl commands: kubectl debug, kubectl events, kubectl top

  • Ephemeral containers for debug running pods

  • Node issues: kubectl describe node, node conditions__HTMLTAG_953___

  • Network debugging with Cilium Hubble

  • _Performance troubleshooting, common issues and solutions

Practice 7:

  • Deploy kube-prometheus-stack (Prometheus + Grafana + AlertManager)

  • Deploy Loki + Grafana Alloy to collect logs_

  • Deploy Tempo and configure OpenTelemetry auto-instrumentation for the application_

  • Create a Grafana dashboard with correlated metrics, logs, traces_

  • Create alerting rule and test notification_


_MODULE 8: ADVANCED TOPICS

Chapter 8.1: Helm 4

  • Helm architecture and Helm 4 vs Helm 3 (released November 2025 — 10th anniversary)

  • Helm 4 new features: WebAssembly (WASM) plugins, server-side apply, 60% performance improvement, OCI enhancements

  • Charts structure, installing and creating custom charts_

  • Chart repositories and OCI registry_

  • Helm hooks, tests, Helmfile_

  • Helm 3 still receives security fixes until November 2026_

Chapter 8.2: Operators and Custom Resources

  • Operator pattern and use cases

  • Custom Resource Definitions (CRDs) and Custom Controllers_

  • Operator SDK and Kubebuilder_

  • Common operators: Prometheus Operator, CloudNativePG (PostgreSQL), Strimzi (Kafka)_

Chapter 8.3: Service Mesh 2026_

  • Why Service Mesh? mTLS, traffic management, observability_

  • Cilium Service Mesh (Sidecarless)_: eBPF at kernel layer, no need for sidecar proxy, 40-60% network reduction overhead

  • Istio: most complete features, suitable for enterprises (multi-cluster, granularity RBAC)

  • Linkerd_: lightest, Rust-based micro-proxy, ideal for resource-constrained environments

  • When to choose what? Comparison chart

Chapter 8.4: GitOps_

  • GitOps principles: Git is single source of truth

  • _ArgoCD 3.x_: centralized, hub-and-spoke multi-cluster, single pane of glass

  • Flux 2.x: decentralized, cluster pulls itself from Git/OCI, safer for distribution teams

  • Choose ArgoCD or Flux? Architectural tradeoffs_

  • CI/CD pipeline with GitHub Actions + ArgoCD/Flux

Practice 8:

  • Create Helm chart with Helm 4, publish to OCI registry

  • Build Simple Operator with Kubebuilder_

  • Compare Cilium Service Mesh vs Istio top cluster

  • Setup GitOps with ArgoCD:deploy application from Git


MODULE 9: CLUSTER MANAGEMENT

Chapter 9.1: Production Cluster Setup

  • kubeadm installation with containerd 2.0 and cgroup v2

  • High Availability: multi-master architecture, load balancing_

  • kube-proxy: nftables mode configuration (IPVS deprecated K8s 1.35)_

  • _Cluster upgrades: safe upgrade strategy for each minor version

  • Backup with Velero: cluster state and PV snapshots_

Chapter 9.2: Infrastructure Migration (important 2026)

  • _Migrate cgroup v1 → cgroup v2: required before upgrading to K8s 1.36

  • Upgrade containerd 1.x → containerd 2.0: required from K8s 1.36

  • Check the compatibility of workloads with cgroup v2_

  • Checklist migration production cluster_

Chapter 9.3: Node Management_

  • Adding/removing nodes

  • Node maintenance: drain, cordon, uncordon_

  • Taints and Tolerations

  • Node affinity, anti-affinity, topology spread constraints_

  • Pod priority and preemption_

Chapter 9.4: Resource Management_

  • Resource requests and limits

  • Quality of Service (QoS) classes: Guaranteed, Burstable, BestEffort_

  • LimitRanges and ResourceQuotas

  • Pod Disruption Budgets (PDB)

  • Cluster Autoscaler vs Karpenter

Chapter 9.5: Cluster API

  • What is Cluster API? Declarative cluster lifecycle management

  • Infrastructure providers: AWS, GCP, Azure, vSphere

  • Create and upgrade clusters with kubectl

Practice 9:

  • Setup 3-node cluster with kubeadm + containerd 2.0 + cgroup v2

  • Perform migrate cgroup v1 → v2 on the existing cluster run_

  • Perform cluster upgrade from 1.32 → 1.33_

  • Install Velero and execute backup/restore

  • Practice disaster recovery scenario


MODULE 10: CLOUD PLATFORMS AND BEST PRACTICES

Chapter 10.1: Managed Kubernetes Services 2026

  • Amazon EKS: Auto Mode, Pod Identity, EKS Anywhere

  • Google GKE: Autopilot, Workload Identity, GKE Enterprise_

  • Azure AKS: Automatic upgrades, KEDA integration, Workload Identity

  • Compare pricing, features, and unique features of each platform_

Chapter 10.2: Cost Optimization

  • Right-sizing with VPA recommendations_

  • Spot/Preemptible for workloads tolerant

  • Karpenter: node consolidation and Spot interruption handling_

  • Kubecost or OpenCost: visibility into per cost namespace/team

Chapter 10.3: Best Practices Production 2026_

  • Cluster setup: multi-AZ, control plane HA

  • Application: resource limits, liveness/readiness probes, PDB_

  • Security hardening according to CIS Kubernetes Benchmark_

  • Multi-tenancy: Namespace isolation, Hierarchical Namespaces Controller_

  • _Hybrid and multi-cloud: Cluster Federation, Multi-cluster Gateway

  • Edge computing with K3s (lightweigh, ARM support)_

Practice 10:

  • Deploy production workload on GKE Autopilot or EKS Auto Mode

  • Install OpenCost and analyze costs by team_

  • Implement Karpenter with Spot instances

  • Final project: deploy complete microservices application with Gateway API, Cilium, GitOps, Observability stack


MODULE 11: AI/ML WORKLOADS ON KUBERNETES (2026)

Chapter 11.1: Kubernetes for AI/ML

  • Why is Kubernetes the ideal platform for AI/ML workloads?

  • GPU support: NVIDIA GPU Operator, installation and configuration_

  • Dynamic Resource Allocation (DRA) GA K8s 1.34_: GPU sharing, FPGA allocation

  • Node selectors and taints/tolerations for GPU nodes_

  • Resource quotas for GPU workloads

Chapter 11.2: Training Jobs_

  • JobSet_: coordinate multiple dependent Jobs in training pipeline

  • Kubeflow Training Operator: PyTorchJob, TFJob, MXJob

  • Distributed training patterns: Data parallelism, Model parallelism

  • Checkpoint and resume training_

Chapter 11.3: Model Serving and Inference

  • Kubernetes Inference Extension (KIE): new standard for LLM serving

  • KServe: model serving framework (TensorFlow, PyTorch, ONNX...)

  • vLLM on Kubernetes for LLM inference_

  • Autoscaling inference with KEDA (scale based on queue depth, GPU utilization)

Chapter 11.4: MLOps Pipelines_

  • Kubeflow Pipelines: orchestrate ML workflows_

  • _Argo Workflows: general-purpose engine_

  • Data processing with Spark on Kubernetes_

  • _Model registry and versioning_

Practice 11:

  • Install NVIDIA GPU Operator and verify GPU scheduling

  • Create PyTorchJob with Kubeflow Training Operator_

  • Deploy LLM inference server with KServe or vLLM_

  • Structure KEDA autoscaling image for inference workload


📖 REFERENCES_

Main Document final

  • Kubernetes Official Documentation: https://kubernetes.io/docs/_

  • Kubernetes Blog: https://kubernetes.io/blog/

  • CNCF Projects: https://www.cncf.io/projects/_

  • Gateway API: https://gateway-api.sigs.k8s.io/

  • OpenTelemetry: https://opentelemetry.io/

Book_

  • "Kubernetes Up & Running" 3rd Ed — Kelsey Hightower (updated 2023)

  • "The Kubernetes Book" — Nigel Poulton (daily update year)

  • "Kubernetes Patterns" 2nd Ed — Bilgin Ibryam & Roland Huß

  • "Production Kubernetes" — Josh Rosso et al.

  • "Cloud Native Observability with OpenTelemetry" — Alex Boten

Courses and Certifications_

  • Certified Kubernetes Administrator (CKA) — Linux Foundation

  • Certified Kubernetes Application Developer (CKAD) — Linux Foundation

  • Certified Kubernetes Security Specialist (CKS) — Linux Foundation

Labs and Practice

  • KillerCoda: https://killercoda.com/ (replaces the closed Katacoda 2023)

  • killer.sh: CKA/CKAD/CKS exam preparation environment_

  • Kubernetes the Hard Way (Kelsey Hightower)_

  • Play with Kubernetes: https://labs.play-with-k8s.com/

Community_

  • Kubernetes Slack: https://slack.k8s.io/

  • Kubernetes GitHub: https://github.com/kubernetes/kubernetes

  • KubeCon + CloudNativeCon conferences


🔧 REQUIRED TOOLS_

Essential Tools

  • kubectl

    ___HTMLTAG_1588__ _
  • kubeadm

  • kind / k3d / Minikube

  • containerd 2.0+ (replace Docker daemon as runtime)

  • VS Code with Kubernetes extension + YAML extension_

Recommended CLI Tools

  • k9s — terminal UI, super fast, "Vim's Kubernetes"

  • kubectx / kubens — switch context/namespace_

  • stern — multi-pod log streaming

  • Helm 4 — package manager

  • kustomize — built-in kubectl, no need to install add

  • cilium CLI — manage and debug Cilium_

Dashboard and IDE

  • Headlamp — official web UI to replace Kubernetes Dashboard (archived January 2026), approved by SIG UI endorse

  • K9s — terminal UI for power users

  • Lens_ — desktop IDE, personal tier free; enterprise tier paid_

  • FreeLens_ — open-source fork of Lens (OpenLens is no longer available maintain)

Observability Stack

  • Prometheus + Grafana + AlertManager

  • Loki (logs) + Tempo (traces) + Grafana Alloy (collector)

  • OpenTelemetry Operator

  • Hubble UI (Cilium network observability)


🎊 CONCLUSION_

Kubernetes in 2026 has matured dramatically: from a container orchestration tool, K8s is becoming into a Universal Control Plane for every workload — containers, VMs, AI/ML pipelines, edge devices. Features like Gateway API, Cilium eBPF, Sidecar containers GA, ValidatingAdmissionPolicy, and Dynamic Resource Allocation show that the ecosystem is getting stronger and more production-ready.

Most important: practice regularly, follow the Kubernetes Blog and CNCF landscape to not be left behind in such a fast-growing ecosystem so.


Note: This course is designed based on Kubernetes version 1.32+ (LTS version suitable for early 2026 production). The latest version at the time of writing is Kubernetes 1.35.3 (March 2026). Check release notes at kubernetes.io/releases before upgrading production cluster.

Module 1: Introduction & Kubernetes Architecture

Module 2: Basic Kubernetes Objects

Module 3: Configuration & Storage

Module 4: Networking

Module 5: Workload Management

Module 6: Security

Module 7: Observability & Monitoring

Module 8: Helm, Operators & GitOps

Module 9: Cluster Management

Module 10: Cloud & Production