
Introduction
Kubernetes (K8s) is the standard orchestration platform for container workloads. Understanding the architecture and core concepts of Kubernetes is mandatory before implementing any microservices system.
1. Kubernetes Architecture
1.1 High-Level Overview
┌────────────────────────────────────────────────────────────┐
│ Kubernetes Cluster │
│ │
│ ┌──────────────────── Control Plane ────────────────────┐ │
│ │ │ │
│ │ ┌───────────┐ ┌──────────┐ ┌───────────────────┐ │ │
│ │ │ API Server│ │Scheduler │ │Controller Manager │ │ │
│ │ │ (kube- │ │ │ │ │ │ │
│ │ │ apiserver)│ │ │ │ - ReplicaSet │ │ │
│ │ └─────┬─────┘ └──────────┘ │ - Deployment │ │ │
│ │ │ │ - Node │ │ │
│ │ ┌─────▼─────┐ │ - Service Account │ │ │
│ │ │ etcd │ └───────────────────┘ │ │
│ │ │ (cluster │ │ │
│ │ │ state) │ ┌──────────────────────────────────┐ │ │
│ │ └───────────┘ │ Cloud Controller Manager │ │ │
│ │ │ (LoadBalancer, Volume, Node) │ │ │
│ │ └──────────────────────────────────┘ │ │
│ └───────────────────────────────────────────────────────┘ │
│ │
│ ┌──────── Worker Node 1 ──────┐ ┌── Worker Node 2 ────┐ │
│ │ │ │ │ │
│ │ ┌────────┐ ┌────────────┐ │ │ ┌────────┐ │ │
│ │ │kubelet │ │kube-proxy │ │ │ │kubelet │ ... │ │
│ │ └────────┘ └────────────┘ │ │ └────────┘ │ │
│ │ ┌────────────────────────┐ │ │ │ │
│ │ │ containerd │ │ │ │ │
│ │ │ ┌─────┐ ┌─────┐ │ │ │ │ │
│ │ │ │Pod A│ │Pod B│ ... │ │ │ │ │
│ │ │ └─────┘ └─────┘ │ │ │ │ │
│ │ └────────────────────────┘ │ │ │ │
│ └──────────────────────────────┘ └──────────────────────┘ │
└────────────────────────────────────────────────────────────┘
1.2 Control Plane Components
kube-apiserver — "Front door" of the cluster:
- Receive all REST API requests
- Authentication (AuthN) and authorization (AuthZ)
- Validate and persist resources into etcd
- Only component that communicates directly with etcd
etcd — Distributed key-value store:
- Stores entire cluster state (desired + actual)
- Strongly consistent (Raft consensus)
- Need to backup regularly
kube-scheduler — Decide which Node the Pod runs on:
- Evaluate resource requests, affinity, taints/tolerations
- Scoring algorithm: choose the most optimal node
kube-controller-manager — Ensure desired state = actual state:
- ReplicaSet Controller: ensures correct number of replicas
- Deployment Controller: manages rolling updates
- Node Controller: detects node failure
- Job Controller: manages one-off tasks
1.3 Worker Node Components
kubelet — Agent running on each node:
- Get Pod spec from API Server
- Make sure the containers in the Pod are running
- Report node status and Pod status
kube-proxy — Network proxy:
- Manage Service networking rules (iptables/IPVS)
- Load balance traffic to Pod endpoints
Container Runtime — containerd or CRI-O:
- Pull images, start/stop containers
- CRI (Container Runtime Interface) compliance
2. Core Resources
2.1 Pod — Smallest unit
Pod is one or more containers running together, sharing network and storage:
apiVersion: v1
kind: Pod
metadata:
name: order-service
labels:
app: order-service
spec:
containers:
- name: order-service
image: registry.example.com/order-service:1.0.0
ports:
- containerPort: 8080
resources:
requests:
cpu: "250m"
memory: "256Mi"
limits:
cpu: "500m"
memory: "512Mi"
env:
- name: DB_HOST
valueFrom:
configMapKeyRef:
name: order-config
key: db_host
Note: In production, never create Pods directly. Always use Deployment.
2.2 Deployment — Lifecycle management
apiVersion: apps/v1
kind: Deployment
metadata:
name: order-service
labels:
app: order-service
spec:
replicas: 3
strategy:
type: RollingUpdate
rollingUpdate:
maxSurge: 1 # Thêm tối đa 1 pod khi update
maxUnavailable: 0 # Không cho phép pod nào unavailable
selector:
matchLabels:
app: order-service
template:
metadata:
labels:
app: order-service
version: v1
spec:
containers:
- name: order-service
image: registry.example.com/order-service:1.0.0
ports:
- containerPort: 8080
resources:
requests:
cpu: "250m"
memory: "512Mi"
limits:
cpu: "500m"
memory: "1Gi"
readinessProbe:
httpGet:
path: /health/ready
port: 8080
initialDelaySeconds: 10
periodSeconds: 5
livenessProbe:
httpGet:
path: /health/live
port: 8080
initialDelaySeconds: 30
periodSeconds: 10
startupProbe:
httpGet:
path: /health/started
port: 8080
failureThreshold: 30
periodSeconds: 2
2.3 Service — Exposure and Load Balance
# ClusterIP — Internal communication (default)
apiVersion: v1
kind: Service
metadata:
name: order-service
spec:
type: ClusterIP
selector:
app: order-service
ports:
- port: 8080
targetPort: 8080
# DNS: order-service.default.svc.cluster.local
# Short: order-service (cùng namespace)
Service types:
ClusterIP → Internal only (default)
NodePort → Expose qua port trên mỗi node (30000-32767)
LoadBalancer → Cloud provider tạo external LB
ExternalName → DNS CNAME alias
2.4 ConfigMap & Secret
# ConfigMap — non-sensitive config
apiVersion: v1
kind: ConfigMap
metadata:
name: order-config
data:
db_host: "postgres-order.database.svc.cluster.local"
db_port: "5432"
db_name: "orders"
log_level: "info"
---
# Secret — sensitive data (base64 encoded)
apiVersion: v1
kind: Secret
metadata:
name: order-secret
type: Opaque
data:
db_password: cGFzc3dvcmQxMjM= # base64
api_key: c2stbXlhcGlrZXk=
Important: Kubernetes Secrets are only base64 encoded, not encoded. In production, use Sealed Secrets or External Secrets Operator + Vault.
2.5 Ingress — External access
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: api-ingress
annotations:
nginx.ingress.kubernetes.io/rate-limit: "100"
cert-manager.io/cluster-issuer: "letsencrypt-prod"
spec:
tls:
- hosts:
- api.example.com
secretName: api-tls
rules:
- host: api.example.com
http:
paths:
- path: /api/orders
pathType: Prefix
backend:
service:
name: order-service
port:
number: 8080
- path: /api/payments
pathType: Prefix
backend:
service:
name: payment-service
port:
number: 8080
3. Auto-Scaling
3.1 Horizontal Pod Autoscaler (HPA)
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: order-service-hpa
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: order-service
minReplicas: 3
maxReplicas: 20
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70
- type: Resource
resource:
name: memory
target:
type: Utilization
averageUtilization: 80
behavior:
scaleDown:
stabilizationWindowSeconds: 300 # Chờ 5 phút trước khi scale down
policies:
- type: Pods
value: 1
periodSeconds: 60 # Giảm tối đa 1 pod mỗi 60s
3.2 Vertical Pod Autoscaler (VPA)
VPA automatically adjusts CPU/Memory requests:
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
name: order-service-vpa
spec:
targetRef:
apiVersion: apps/v1
kind: Deployment
name: order-service
updatePolicy:
updateMode: "Auto"
resourcePolicy:
containerPolicies:
- containerName: order-service
minAllowed:
cpu: "100m"
memory: "128Mi"
maxAllowed:
cpu: "2"
memory: "4Gi"
4. Namespace Strategy
Kubernetes Cluster
│
├── kube-system # System components (CoreDNS, metrics-server)
├── kube-public # Public resources
│
├── platform # Shared infrastructure
│ ├── kafka
│ ├── redis
│ ├── postgresql
│ └── prometheus
│
├── gateway # API Gateway (Kong/Traefik)
│
├── services-prod # Production services
│ ├── order-service (3 replicas)
│ ├── payment-service (3 replicas)
│ ├── inventory-service (2 replicas)
│ └── notification-service (2 replicas)
│
├── services-staging # Staging (1 replica each)
│
├── monitoring # Observability stack
│ ├── grafana
│ ├── loki
│ ├── jaeger
│ └── alertmanager
│
└── argocd # GitOps controller
Resource Quotas per namespace:
apiVersion: v1
kind: ResourceQuota
metadata:
name: services-prod-quota
namespace: services-prod
spec:
hard:
requests.cpu: "20"
requests.memory: "40Gi"
limits.cpu: "40"
limits.memory: "80Gi"
pods: "100"
5. Kubernetes Networking Model
5.1 Pod-to-Pod Communication
Kubernetes Network Rules:
1. Mọi Pod có thể giao tiếp với mọi Pod khác (không cần NAT)
2. Agents trên node (kubelet) có thể giao tiếp với tất cả Pods trên node đó
3. Mỗi Pod có IP riêng trong cluster CIDR
Pod A (10.244.1.5) ──────────▶ Pod B (10.244.2.3)
Node 1 Node 2
│ │
└──── CNI Plugin (Calico/Cilium/Flannel) ────┘
5.2 Service Discovery via CoreDNS
Service DNS formats:
├── <service>.<namespace>.svc.cluster.local (FQDN)
├── <service>.<namespace>.svc (shortened)
├── <service>.<namespace> (shortened)
└── <service> (same namespace)
Ví dụ:
order-service gọi payment-service:
curl http://payment-service:8080/api/pay # same namespace
curl http://payment-service.services-prod:8080 # cross namespace
6. Summary
| Components | Role |
|---|---|
| Control Plane | Manage cluster state, scheduling, controller loops |
| Pod | Smallest deployment unit, 1+ containers |
| Deployment | Manage ReplicaSet, rolling update, rollback |
| Service | Stable network endpoint, load balancing |
| ConfigMap/Secret | Externalize configuration |
| Ingress | External HTTP(S) routing |
| HPA/VPA | Auto-scaling horizontal and vertical |
| Namespace | Logical isolation, resource quotas |
Next post: Microservices Design Principles — SRP, DDD & Bounded Context to properly define service boundaries.