🎯 LESSON OBJECTIVE__HTMLTAG_68___
After completing this lesson, you will:
- ✅ Understand the difference between on-premises, cloud, and hybrid deployments for microservices
- ✅ Understand the architectural overview and all core components of the production system
- ✅ Understand the reasons for choosing each technology in the stack (Kubernetes, Ceph, Patroni, Istio, ArgoCD...)
- ✅ Set up the lab environment for the entire course
- ✅ Understand the roadmap of 50 lessons and the links between sections__HTMLTAG_81___
PART 1: WHY ON-PREMISES FOR MICROSERVICES?
1.1. Actual context
In the cloud-native era, many organizations still choose to deploy on-premises because:
📊 Actual statistics (2025-2026):
- ~60% of enterprise workloads still run on-premises or hybrid (Gartner)
- Cloud cost increases 30-40% each year when scaling → "cloud repatriation" trend
- Regulated industries (finance, healthcare, government) require data sovereignty__HTMLTAG_100___
- Latency-sensitive applications need proximity to users/devices
1.2. Compare On-Premises vs Cloud vs Hybrid
| Criteria | On-Premises | Public Cloud | Hybrid |
|---|---|---|---|
| Initial Cost (CapEx) | High (buy hardware) | Low (pay-as-you-go) | Average |
| Long-Term Expenses (OpEx) | Lower when scaled__HTMLTAG_139___ | High and unpredictable__HTMLTAG_141___ | Depending on workload |
| Data Sovereignty | ✅ Full control__HTMLTAG_151___ | ⚠️ Region dependent | ✅ Mostly on-prem |
| Latency | ✅ Lowest | Region Dependent | Good for edge cases |
| Customization | ✅ Unlimited | Limited by provider | Flexible__HTMLTAG_179___ |
| Ops Complexity | ❌ High (self-managed) | ✅ Low (managed) | Highest |
| Scaling Speed | ❌ Slow (buying hardware) | ✅ Minutes (auto-scale) | Flexible |
| Compliance | ✅ Easiest to respond | Need shared responsibility | Good |
| Vendor Lock-in | ✅ No | ❌ High (AWS/GCP/Azure) | Average |
1.3. When should you choose On-Premises?
✅ Should choose On-Premises when:
- Workloads are stable, predictable (not bursting up and down continuously)
- High compliance requirements (HIPAA, PCI-DSS, GDPR data residency)
- Infrastructure investment (data center, servers, networking)
- Cloud monthly cost exceeds threshold (~$50K-100K+/month)
- DevOps/SRE Team with operational experience__HTMLTAG_248___
- Needs ultra-low latency (< 1ms giữa services)
❌ Do not select On-Premises when:
- Early-stage startups need speed to market
- Workloads bursty, difficult to predict__HTMLTAG_260___
- Team < 5 người, không có infra engineer
- PoC/MVP needs to be deployed quickly__HTMLTAG_264___
PART 2: OVERALL SYSTEM ARCHITECTURE
2.1. Overall architecture diagram
graph TB
subgraph EA["🌐 EXTERNAL ACCESS"]
Users["👤 Users"] --> DNS["DNS"]
DNS --> MetalLB["MetalLB VIP"]
MetalLB --> NGINX["NGINX Ingress"]
NGINX --> Gateway["Istio Gateway
+ cert-manager TLS"]
end
subgraph K8S["☸ KUBERNETES HA CLUSTER — 3 Control Plane + N Workers"]
subgraph MESH["🔒 Service Mesh — Istio"]
mTLS["mTLS"] ~~~ TM["Traffic Mgmt"] ~~~ CB["Circuit Breaker"] ~~~ CD["Canary Deploy"]
end
subgraph MS["📦 Microservices"]
APIGW["API Gateway"]
AuthSvc["Auth Service"]
UserSvc["User Service"]
OrderSvc["Order Service"]
PaySvc["Payment Service"]
NotifSvc["Notification Service"]
end
subgraph DL["💾 Data Layer"]
PG["PostgreSQL HA
CloudNativePG + PgBouncer"]
Redis["Redis HA
Sentinel / Cluster"]
RMQ["RabbitMQ HA
Quorum Queues"]
Kafka["Kafka
Strimzi KRaft"]
end
subgraph GO["🔄 GitOps & Secrets"]
ArgoCD["ArgoCD HA"]
Helm["Helm Charts"]
Vault["Vault HA + ESO"]
Kyverno["Kyverno Policies"]
end
subgraph OBS["📊 Observability"]
Prom["Prometheus HA + Thanos"]
Grafana["Grafana HA"]
Loki["Loki + Alloy"]
Tempo["Tempo + OTEL"]
end
subgraph SEC["🛡️ Security"]
RBAC["RBAC + OIDC
Keycloak"]
Falco["Falco Runtime"]
Harbor["Trivy + Harbor"]
NP["NetworkPolicy
Cilium"]
end
subgraph STOR["💿 Storage — Rook-Ceph"]
RBD["RBD Block
→ Databases"]
CephFS["CephFS Shared
→ Apps"]
RGW["RGW / S3 Object
→ Backup"]
end
subgraph INFRA["⚙️ Infrastructure"]
CiliumCNI["Cilium CNI eBPF"]
MetalLBi["MetalLB"]
CoreDNS["CoreDNS"]
etcd["etcd HA"]
HAVIP["keepalived + HAProxy
API Server VIP"]
end
end
subgraph PHYS["🖥️ Physical Layer"]
CP["3× Control Plane Nodes"] ~~~ WK["3-5× Worker Nodes"] ~~~ SN["3× Storage Nodes"]
NET["Network: Mgmt + Cluster + Storage + External VLANs
OS: Ubuntu 24.04 LTS / RHEL 9"]
end
Gateway --> MESH
MESH --> MS
MS --> DL
K8S --> PHYS
2.2. Core components and roles
Layer 1: Infrastructure Foundation
| Element__HTMLTAG_280___ | Technology | Role | Lesson__HTMLTAG_286___ |
|---|---|---|---|
| Container Runtime | containerd 2.x | Run containers according to CRI standard | Lesson 5 |
| K8s Orchestration__HTMLTAG_302___ | kubeadm (K8s 1.31+) | HA control plane, scheduling, self-healing__HTMLTAG_306___ | Lesson 5-7 |
| CNI Networking__HTMLTAG_312___ | Cilium (eBPF) | Pod networking, NetworkPolicy, Hubble observability__HTMLTAG_316___ | Lesson 8 |
| Load Balancer | MetalLB | Grant External IP to Services on bare-metal | Lesson 9 |
| API Server HA | keepalived + HAProxy | Virtual IP for K8s API endpoint__HTMLTAG_336___ | Lesson 4 |
| Cluster State | etcd (3 nodes) | Distributed key-value store for K8s | Lesson 10 |
Layer 2: Distributed Storage
| Element__HTMLTAG_360___ | Technology | Role | Lesson |
|---|---|---|---|
| Storage Orchestrator | Rook Operator | Manage Ceph lifecycle on K8s | Lesson 11-12 |
| Block Storage | Ceph RBD | PV for databases (PostgreSQL, etcd) | Lesson 13 |
| Shared Storage | CephFS | ReadWriteMany for microservices__HTMLTAG_396___ | Lesson 14 |
| Object Storage | Ceph RGW (S3) | Backup, Loki logs, Thanos metrics | Lesson 15 |
Layer 3: Data Layer
| Element | Technology | Role | Lesson__HTMLTAG_426___ |
|---|---|---|---|
| Primary Database | PostgreSQL HA (CloudNativePG) | ACID transactions, relational data | Lessons 16-17 |
| Connection Pool | PgBouncer | Connection pooling, reduce DB load__HTMLTAG_446___ | Lesson 18 |
| DB Backup | pgBackRest | Full/incremental backup, PITR | Lesson 19 |
| Message Queue | RabbitMQ HA | Async messaging, task queues__HTMLTAG_466___ | Lesson 21 |
| Event Streaming__HTMLTAG_472___ | Kafka (Strimzi) | Event sourcing, log aggregation | Lesson 22 |
| Cache | Redis HA | Caching, session store, rate limiting__HTMLTAG_486___ | Lesson 23 |
Layer 4: Service Mesh & Networking
| Element | Technology | Role | Lesson__HTMLTAG_506___ |
|---|---|---|---|
| Service Mesh | Istio | mTLS, traffic management, observability__HTMLTAG_516___ | Lessons 24-25 |
| Ingress Controller | NGINX Ingress | HTTP/HTTPS routing into cluster | Lesson 26 |
| TLS Automation | cert-manager | Auto-issue/renew certificates | Lesson 26 |
| Gateway API | Istio + Gateway API | Next-gen ingress, canary routing | Lesson 27 |
Layer 5: Platform Operations
| Element__HTMLTAG_560___ | Technology | Role | Lesson |
|---|---|---|---|
| GitOps | ArgoCD HA | Declarative deployment from Git | Lessons 28, 30 |
| Packaging | Helm | K8s manifest templating | Lesson 29 |
| Secrets | Vault HA + ESO | Centralized secrets management | Lesson 31 |
| Metrics | Prometheus HA + Thanos | Metrics collection, long-term storage | Lesson 32 |
| Dashboards | Grafana HA | Visualization, alerting | Lesson 33 |
| Logs | Loki + Alloy | Centralized log aggregation | Lesson 34 |
| Traces | Tempo + OpenTelemetry__HTMLTAG_634___ | Distributed tracing | Lesson 35 |
| Policy | Kyverno | Admission control, policy-as-code | Lesson 37 |
| Runtime Security | Falco | Threat detection | Lesson 38 |
| Image Security | Trivy + Harbor | Vulnerability scanning, private registry | Lesson 39 |
| Backup | Velero | Cluster backup/restore | Lesson 44 |
| Chaos Testing | Chaos Mesh | Resilience validation | Lesson 45 |
PART 3: WHY CHOOSE EACH TECHNOLOGY?
3.1. Kubernetes (kubeadm) — Why not use managed K8s?
On-premises does not have EKS/GKE/AKS. Options:
| Tool | Advantages | Disadvantages | Relevant |
|---|---|---|---|
| kubeadm | Official K8s tool, flexible, production-grade | Manual setup, need to understand deeply | ✅ Production |
| k3s | Lightweight, easy to install__HTMLTAG_731___ | Remove features, use SQLite instead etcd | Edge/IoT |
| RKE2 | FIPS compliant, Rancher integration | Vendor-specific | Rancher users |
| Kubespray | Ansible-based, reproducible | Slow, Ansible complexity | Large clusters |
👉 Choose kubeadm because: official tool, production-grade, helps understand K8s internals the deepest.
3.2. Cilium CNI — Why not Calico or Flannel?
Flannel: Đơn giản → Không có NetworkPolicy → ❌ Production
Calico: Tốt → iptables-based → Performance overhead khi scale
Cilium: eBPF-based → Kernel-level networking → ✅ Best performance
+ Hubble observability + kube-proxy replacement
+ CNCF Graduated project (2024)
3.3. Rook-Ceph — Why not Longhorn or NFS?
NFS: Single point of failure, no replication → ❌ HA
Longhorn: Đơn giản, tốt cho small clusters → Không có Object Storage
Rook-Ceph: Block + Shared + Object storage trong 1 platform
Enterprise-grade, CNCF Graduated
Performance tốt cho databases + S3 cho backup/logs
→ ✅ All-in-one storage solution
3.4. Istio — Why not Linkerd?
Linkerd: Nhẹ hơn, dễ hơn → Ít features (không Gateway API, limited traffic mgmt)
Istio: Feature-rich → mTLS, traffic mirroring, canary, circuit breaker
Gateway API support, Kiali observability
Industry standard cho enterprise → ✅ Production choice
PART 4: ENVIRONMENT LAB SETUP
4.1. Minimum Hardware for Lab
You need at least the following resources to practice the entire course:
Option A: VMs on powerful hosts (Recommended)
block-beta
columns 3
block:HOST["🖥️ Host Machine: 64GB RAM, 16 cores, 500GB SSD"]:3
block:CP["Control Plane Nodes"]:1
m1["master1
4 vCPU · 8GB RAM
50GB disk"]
m2["master2
4 vCPU · 8GB RAM
50GB disk"]
m3["master3
4 vCPU · 8GB RAM
50GB disk"]
end
block:WK["Worker Nodes"]:1
w1["worker1
4 vCPU · 8GB RAM
50GB + 100GB raw"]
w2["worker2
4 vCPU · 8GB RAM
50GB + 100GB raw"]
w3["worker3
4 vCPU · 8GB RAM
50GB + 100GB raw"]
end
block:LB["Load Balancer"]:1
lb["lb
2 vCPU · 2GB RAM
20GB disk
HAProxy + keepalived"]
end
end
style CP fill:#1e3a5f,stroke:#3b82f6,color:#e2e8f0
style WK fill:#1e3a5f,stroke:#10b981,color:#e2e8f0
style LB fill:#1e3a5f,stroke:#f59e0b,color:#e2e8f0
Total: ~26 vCPU, 58GB RAM, 520GB disk
Option B: Cloud VMs (AWS/GCP/Hetzner)
7 VMs tương đương cấu hình bên trên
Estimated cost: ~$200-400/tháng (Hetzner rẻ nhất)
Khuyến nghị: Hetzner Dedicated hoặc Proxmox VE
Option C: Bare-metal (Production-like)
3× Dell PowerEdge R640 hoặc tương đương:
- 2× 16-core Xeon, 128GB RAM, 2× 480GB SSD (OS) + 4× 2TB NVMe (Ceph)
- 4× 25GbE NICs (bonding)
4.2. Network Layout for Lab
graph TB
subgraph MGMT["🌐 Management Network — 192.168.1.0/24"]
direction LR
lb["lb
192.168.1.10"]
m1["master1
192.168.1.11"]
m2["master2
192.168.1.12"]
m3["master3
192.168.1.13"]
w1["worker1
192.168.1.21"]
w2["worker2
192.168.1.22"]
w3["worker3
192.168.1.23"]
VIP["🔷 VIP
192.168.1.100
K8s API Server"]
end
subgraph INTERNAL["🔒 Internal Networks"]
POD["Pod Network
10.244.0.0/16
Cilium CNI"]
SVC["Service Network
10.96.0.0/12
ClusterIP"]
LB_POOL["MetalLB Pool
192.168.1.200–250
External Services"]
end
VIP --> m1 & m2 & m3
lb --> VIP
style VIP fill:#dc2626,stroke:#fca5a5,color:#fff
style POD fill:#1e3a5f,stroke:#3b82f6,color:#e2e8f0
style SVC fill:#1e3a5f,stroke:#10b981,color:#e2e8f0
style LB_POOL fill:#1e3a5f,stroke:#f59e0b,color:#e2e8f0
4.3. Create VMs Fast with Vagrant (Optional)
# Vagrantfile Vagrant.configure("2") do |config| config.vm.box = "ubuntu/noble64" # Ubuntu 24.04Load Balancer
config.vm.define "lb" do |lb| lb.vm.hostname = "lb" lb.vm.network "private_network", ip: "192.168.1.10" lb.vm.provider "virtualbox" do |v| v.memory = 2048 v.cpus = 2 end end
Control Plane nodes
(1..3).each do |i| config.vm.define "master#{i}" do |master| master.vm.hostname = "master#{i}" master.vm.network "private_network", ip: "192.168.1.#{10 + i}" master.vm.provider "virtualbox" do |v| v.memory = 8192 v.cpus = 4 end end end
Worker nodes
(1..3).each do |i| config.vm.define "worker#{i}" do |worker| worker.vm.hostname = "worker#{i}" worker.vm.network "private_network", ip: "192.168.1.#{20 + i}" worker.vm.provider "virtualbox" do |v| v.memory = 8192 v.cpus = 4 # Raw disk cho Ceph OSD unless File.exist?("ceph-osd-worker#{i}.vdi") v.customize ['createmedium', 'disk', '--filename', "ceph-osd-worker#{i}.vdi", '--size', 102400] end v.customize ['storageattach', :id, '--storagectl', 'SCSI', '--port', 2, '--type', 'hdd', '--medium', "ceph-osd-worker#{i}.vdi"] end end end end
# Khởi tạo toàn bộ lab
vagrant up
# SSH vào master1
vagrant ssh master1
# Kiểm tra connectivity
for i in 10 11 12 13 21 22 23; do
ping -c 1 192.168.1.$i
done
4.4. Configure SSH Keys for all nodes
# Trên máy workstation/jump host ssh-keygen -t ed25519 -C "k8s-lab-admin" -f ~/.ssh/k8s-labCopy public key sang tất cả nodes
for host in lb master{1..3} worker{1..3}; do ssh-copy-id -i ~/.ssh/k8s-lab.pub user@${host} done
Tạo SSH config cho tiện
cat >> ~/.ssh/config << 'EOF' Host lb HostName 192.168.1.10 User root
Host master1 HostName 192.168.1.11 User root
Host master2 HostName 192.168.1.12 User root
Host master3 HostName 192.168.1.13 User root
Host worker1 HostName 192.168.1.21 User root
Host worker2 HostName 192.168.1.22 User root
Host worker3 HostName 192.168.1.23 User root
Host master* worker* lb IdentityFile ~/.ssh/k8s-lab StrictHostKeyChecking no EOF
PART 5: LEARNING ROUTE 50 LESSON
5.1. Dependency Graph between sections
graph TD
P1["📐 Phase 1: Foundation
Bài 1-4
Chuẩn bị hạ tầng cơ bản"]
P2["☸ Phase 2: K8s HA
Bài 5-10
Dựng Kubernetes HA cluster"]
P3["💿 Phase 3: Rook-Ceph
Bài 11-15
Distributed Storage"]
P4["🐘 Phase 4: PostgreSQL HA
Bài 16-20"]
P5["📨 Phase 5: MQ HA
Bài 21-23
RabbitMQ · Kafka · Redis"]
P6["🔗 Phase 6: Istio
Bài 24-27
Service Mesh"]
P7["🔄 Phase 7: GitOps
Bài 28-31
ArgoCD + Helm + Vault"]
P8["📊 Phase 8: Observability
Bài 32-35"]
P9["🛡️ Phase 9: Security
Bài 36-39"]
P10["🚀 Phase 10: Deployment Patterns
Bài 40-43"]
P11["💥 Phase 11: DR & Chaos
Bài 44-45"]
P12["🏭 Phase 12: Operations + Capstone
Bài 46-50"]
P1 --> P2
P2 --> P3 & P5 & P6
P3 --> P4
P4 & P5 & P6 --> P7
P7 --> P8 & P9
P8 & P9 --> P10
P10 --> P11
P11 --> P12
style P1 fill:#1e40af,stroke:#3b82f6,color:#fff
style P2 fill:#1e40af,stroke:#3b82f6,color:#fff
style P12 fill:#15803d,stroke:#22c55e,color:#fff
5.2. Estimated time
| Section | Post number | Time | Timeline (2 hours/day) |
|---|---|---|---|
| Part 1: Foundation | 4 | ~8h | Week 1 |
| Part 2: K8s HA | 6 | ~14h | Weeks 2-3 |
| Part 3: Rook-Ceph | 5 | ~11h | Week 3-4 |
| Part 4: PostgreSQL__HTMLTAG_847___ | 5 | ~12h | Week 5-6 |
| Part 5: MQ HA | 3 | ~8h | Week 6-7 |
| Part 6: Istio | 4 | ~10h | Week 7-8 |
| Part 7: GitOps | 4 | ~11h | Week 9-10 |
| Part 8: Observability__HTMLTAG_887___ | 4 | ~10h | Week 10-11 |
| Part 9: Security | 4 | ~10h | Week 12-13 |
| Part 10: Deployment | 4 | ~9h | Week 13-14 |
| Part 11: DR | 2 | ~5h | Week 15 |
| Part 12: Operations | 5 | ~15h | Week 15-18 |
| TOTAL | 50 | ~123h | ~18 weeks |
PART 6: CONVENTIONS AND CONVENTIONS IN THE COURSE
6.1. Naming Conventions
# Namespace naming production: prod-<service-name> # prod-user-service staging: stg-<service-name> infrastructure: infra-<component> # infra-monitoring, infra-storage platform: platform-<component> # platform-argocd, platform-vaultHelm release naming
<component>-<environment> # postgresql-prod, redis-stg
Label standards
app.kubernetes.io/name: <service-name> app.kubernetes.io/version: <version> app.kubernetes.io/component: <component> app.kubernetes.io/part-of: <system-name> app.kubernetes.io/managed-by: helm
6.2. Symbols in the lesson
- 💡 Tip: Useful tips, best practice
- ⚠️ Warning: Be careful, it may cause errors
- ❌ Danger: Absolutely do not do it in production
- 📋 Checklist: Checklist
- 🔬 Deep Dive: Technical explanation
- 🛠️ Lab: Practice
💡 KEY TAKEAWAYS
- On-premises microservices suitable for organizations that need data sovereignty, predictable cost, and ultra-low latency
- Kubernetes HA is an orchestration platform, combined with the CNCF tools ecosystem to create a production platform
- Full stack__HTMLTAG_1003___ includes 6 layers: Infrastructure → Storage → Data → Networking → Platform → Security
- Each technology is selectedbased on criteria: production-grade, CNCF backed, community active
- Lab environment requires a minimum of 7 VMs (3 masters + 3 workers + 1 LB) with ~58GB total RAM__HTMLTAG_1012___
🎯 EXERCISE
Exercise 1: Assessing infrastructure requirements__HTMLTAG_1018___
For scenario: Fintech company needs to deploy 20 microservices, handle 10,000 requests/second, store 500GB of data, require PCI-DSS compliance.
- Calculate the number of nodes needed (control plane, workers, storage)
- Estimated total CPU, RAM, Storage
- Draw network topology diagram
- List the required components from the above stack
Exercise 2: Setup Lab Environment__HTMLTAG_1032___
- Create 7 VMs using Option A or Option B
- Configuring networking between VMs
- Setup SSH key-based authentication
- Verify ping connectivity between all nodes
- Record the IP and hostname of each VM
Exercise 3: Comparing technology__HTMLTAG_1046___
Research and compare in detail two pairs of technologies:
- Cilium vs Calico: Performance benchmarks, features, community
- Rook-Ceph vs Longhorn: Scalability, features, operational complexity
📚 NEXT POST
In Lesson 2: Hardware Planning and Network Topology, we will dive into detailed sizing calculations for CPU/RAM/Disk, network topology design with VLAN, bonding, and MTU for production environment.