Chuyển đến nội dung chính

LESSON 1: OVERVIEW OF MICROSERVICES ON-PREMISES ARCHITECTURE

Compare on-premises vs cloud vs hybrid, core components of a microservices production system (K8s, DB HA, Storage, Messaging, Observability, Security), learning path and lab environment setup.

🔒 DevSecOps — Lesson 1 LESSON 1: MICROSERVICES ARCHITECTURE OVERVIEW ON-PREMISES

Deploy Microservices On-Premises with Kubernetes HA

Part 1: Platform & On-Premises Infrastructure Design

xdev.asia

🎯 LESSON OBJECTIVE__HTMLTAG_68___

After completing this lesson, you will:

  • ✅ Understand the difference between on-premises, cloud, and hybrid deployments for microservices
  • ✅ Understand the architectural overview and all core components of the production system
  • ✅ Understand the reasons for choosing each technology in the stack (Kubernetes, Ceph, Patroni, Istio, ArgoCD...)
  • ✅ Set up the lab environment for the entire course
  • ✅ Understand the roadmap of 50 lessons and the links between sections__HTMLTAG_81___

PART 1: WHY ON-PREMISES FOR MICROSERVICES?

1.1. Actual context

In the cloud-native era, many organizations still choose to deploy on-premises because:

📊 Actual statistics (2025-2026):

  • ~60% of enterprise workloads still run on-premises or hybrid (Gartner)
  • Cloud cost increases 30-40% each year when scaling → "cloud repatriation" trend
  • Regulated industries (finance, healthcare, government) require data sovereignty__HTMLTAG_100___
  • Latency-sensitive applications need proximity to users/devices

1.2. Compare On-Premises vs Cloud vs Hybrid

Criteria On-Premises Public Cloud Hybrid
Initial Cost (CapEx) High (buy hardware) Low (pay-as-you-go) Average
Long-Term Expenses (OpEx) Lower when scaled__HTMLTAG_139___ High and unpredictable__HTMLTAG_141___ Depending on workload
Data Sovereignty ✅ Full control__HTMLTAG_151___ ⚠️ Region dependent ✅ Mostly on-prem
Latency ✅ Lowest Region Dependent Good for edge cases
Customization ✅ Unlimited Limited by provider Flexible__HTMLTAG_179___
Ops Complexity ❌ High (self-managed) ✅ Low (managed) Highest
Scaling Speed ❌ Slow (buying hardware) ✅ Minutes (auto-scale) Flexible
Compliance ✅ Easiest to respond Need shared responsibility Good
Vendor Lock-in ✅ No ❌ High (AWS/GCP/Azure) Average

1.3. When should you choose On-Premises?

✅ Should choose On-Premises when:

  • Workloads are stable, predictable (not bursting up and down continuously)
  • High compliance requirements (HIPAA, PCI-DSS, GDPR data residency)
  • Infrastructure investment (data center, servers, networking)
  • Cloud monthly cost exceeds threshold (~$50K-100K+/month)
  • DevOps/SRE Team with operational experience__HTMLTAG_248___
  • Needs ultra-low latency (< 1ms giữa services)

❌ Do not select On-Premises when:

  • Early-stage startups need speed to market
  • Workloads bursty, difficult to predict__HTMLTAG_260___
  • Team < 5 người, không có infra engineer
  • PoC/MVP needs to be deployed quickly__HTMLTAG_264___

PART 2: OVERALL SYSTEM ARCHITECTURE

2.1. Overall architecture diagram


graph TB
    subgraph EA["🌐 EXTERNAL ACCESS"]
        Users["👤 Users"] --> DNS["DNS"]
        DNS --> MetalLB["MetalLB VIP"]
        MetalLB --> NGINX["NGINX Ingress"]
        NGINX --> Gateway["Istio Gateway
+ cert-manager TLS"] end subgraph K8S["☸ KUBERNETES HA CLUSTER — 3 Control Plane + N Workers"] subgraph MESH["🔒 Service Mesh — Istio"] mTLS["mTLS"] ~~~ TM["Traffic Mgmt"] ~~~ CB["Circuit Breaker"] ~~~ CD["Canary Deploy"] end subgraph MS["📦 Microservices"] APIGW["API Gateway"] AuthSvc["Auth Service"] UserSvc["User Service"] OrderSvc["Order Service"] PaySvc["Payment Service"] NotifSvc["Notification Service"] end subgraph DL["💾 Data Layer"] PG["PostgreSQL HA
CloudNativePG + PgBouncer"] Redis["Redis HA
Sentinel / Cluster"] RMQ["RabbitMQ HA
Quorum Queues"] Kafka["Kafka
Strimzi KRaft"] end subgraph GO["🔄 GitOps & Secrets"] ArgoCD["ArgoCD HA"] Helm["Helm Charts"] Vault["Vault HA + ESO"] Kyverno["Kyverno Policies"] end subgraph OBS["📊 Observability"] Prom["Prometheus HA + Thanos"] Grafana["Grafana HA"] Loki["Loki + Alloy"] Tempo["Tempo + OTEL"] end subgraph SEC["🛡️ Security"] RBAC["RBAC + OIDC
Keycloak"] Falco["Falco Runtime"] Harbor["Trivy + Harbor"] NP["NetworkPolicy
Cilium"] end subgraph STOR["💿 Storage — Rook-Ceph"] RBD["RBD Block
→ Databases"] CephFS["CephFS Shared
→ Apps"] RGW["RGW / S3 Object
→ Backup"] end subgraph INFRA["⚙️ Infrastructure"] CiliumCNI["Cilium CNI eBPF"] MetalLBi["MetalLB"] CoreDNS["CoreDNS"] etcd["etcd HA"] HAVIP["keepalived + HAProxy
API Server VIP"] end end subgraph PHYS["🖥️ Physical Layer"] CP["3× Control Plane Nodes"] ~~~ WK["3-5× Worker Nodes"] ~~~ SN["3× Storage Nodes"] NET["Network: Mgmt + Cluster + Storage + External VLANs
OS: Ubuntu 24.04 LTS / RHEL 9"] end Gateway --> MESH MESH --> MS MS --> DL K8S --> PHYS

2.2. Core components and roles

Layer 1: Infrastructure Foundation

Element__HTMLTAG_280___ Technology Role Lesson__HTMLTAG_286___
Container Runtime containerd 2.x Run containers according to CRI standard Lesson 5
K8s Orchestration__HTMLTAG_302___ kubeadm (K8s 1.31+) HA control plane, scheduling, self-healing__HTMLTAG_306___ Lesson 5-7
CNI Networking__HTMLTAG_312___ Cilium (eBPF) Pod networking, NetworkPolicy, Hubble observability__HTMLTAG_316___ Lesson 8
Load Balancer MetalLB Grant External IP to Services on bare-metal Lesson 9
API Server HA keepalived + HAProxy Virtual IP for K8s API endpoint__HTMLTAG_336___ Lesson 4
Cluster State etcd (3 nodes) Distributed key-value store for K8s Lesson 10

Layer 2: Distributed Storage

Element__HTMLTAG_360___ Technology Role Lesson
Storage Orchestrator Rook Operator Manage Ceph lifecycle on K8s Lesson 11-12
Block Storage Ceph RBD PV for databases (PostgreSQL, etcd) Lesson 13
Shared Storage CephFS ReadWriteMany for microservices__HTMLTAG_396___ Lesson 14
Object Storage Ceph RGW (S3) Backup, Loki logs, Thanos metrics Lesson 15

Layer 3: Data Layer

Element Technology Role Lesson__HTMLTAG_426___
Primary Database PostgreSQL HA (CloudNativePG) ACID transactions, relational data Lessons 16-17
Connection Pool PgBouncer Connection pooling, reduce DB load__HTMLTAG_446___ Lesson 18
DB Backup pgBackRest Full/incremental backup, PITR Lesson 19
Message Queue RabbitMQ HA Async messaging, task queues__HTMLTAG_466___ Lesson 21
Event Streaming__HTMLTAG_472___ Kafka (Strimzi) Event sourcing, log aggregation Lesson 22
Cache Redis HA Caching, session store, rate limiting__HTMLTAG_486___ Lesson 23

Layer 4: Service Mesh & Networking

Element Technology Role Lesson__HTMLTAG_506___
Service Mesh Istio mTLS, traffic management, observability__HTMLTAG_516___ Lessons 24-25
Ingress Controller NGINX Ingress HTTP/HTTPS routing into cluster Lesson 26
TLS Automation cert-manager Auto-issue/renew certificates Lesson 26
Gateway API Istio + Gateway API Next-gen ingress, canary routing Lesson 27

Layer 5: Platform Operations

Element__HTMLTAG_560___ Technology Role Lesson
GitOps ArgoCD HA Declarative deployment from Git Lessons 28, 30
Packaging Helm K8s manifest templating Lesson 29
Secrets Vault HA + ESO Centralized secrets management Lesson 31
Metrics Prometheus HA + Thanos Metrics collection, long-term storage Lesson 32
Dashboards Grafana HA Visualization, alerting Lesson 33
Logs Loki + Alloy Centralized log aggregation Lesson 34
Traces Tempo + OpenTelemetry__HTMLTAG_634___ Distributed tracing Lesson 35
Policy Kyverno Admission control, policy-as-code Lesson 37
Runtime Security Falco Threat detection Lesson 38
Image Security Trivy + Harbor Vulnerability scanning, private registry Lesson 39
Backup Velero Cluster backup/restore Lesson 44
Chaos Testing Chaos Mesh Resilience validation Lesson 45

PART 3: WHY CHOOSE EACH TECHNOLOGY?

3.1. Kubernetes (kubeadm) — Why not use managed K8s?

On-premises does not have EKS/GKE/AKS. Options:

Tool Advantages Disadvantages Relevant
kubeadm Official K8s tool, flexible, production-grade Manual setup, need to understand deeply ✅ Production
k3s Lightweight, easy to install__HTMLTAG_731___ Remove features, use SQLite instead etcd Edge/IoT
RKE2 FIPS compliant, Rancher integration Vendor-specific Rancher users
Kubespray Ansible-based, reproducible Slow, Ansible complexity Large clusters

👉 Choose kubeadm because: official tool, production-grade, helps understand K8s internals the deepest.

3.2. Cilium CNI — Why not Calico or Flannel?


Flannel:  Đơn giản → Không có NetworkPolicy → ❌ Production
Calico:   Tốt → iptables-based → Performance overhead khi scale
Cilium:   eBPF-based → Kernel-level networking → ✅ Best performance
          + Hubble observability + kube-proxy replacement
          + CNCF Graduated project (2024)

3.3. Rook-Ceph — Why not Longhorn or NFS?


NFS:      Single point of failure, no replication → ❌ HA
Longhorn: Đơn giản, tốt cho small clusters → Không có Object Storage
Rook-Ceph: Block + Shared + Object storage trong 1 platform
           Enterprise-grade, CNCF Graduated
           Performance tốt cho databases + S3 cho backup/logs
           → ✅ All-in-one storage solution

3.4. Istio — Why not Linkerd?


Linkerd: Nhẹ hơn, dễ hơn → Ít features (không Gateway API, limited traffic mgmt)
Istio:   Feature-rich → mTLS, traffic mirroring, canary, circuit breaker
         Gateway API support, Kiali observability
         Industry standard cho enterprise → ✅ Production choice

PART 4: ENVIRONMENT LAB SETUP

4.1. Minimum Hardware for Lab

You need at least the following resources to practice the entire course:

Option A: VMs on powerful hosts (Recommended)


block-beta
    columns 3
    block:HOST["🖥️ Host Machine: 64GB RAM, 16 cores, 500GB SSD"]:3
        block:CP["Control Plane Nodes"]:1
            m1["master1
4 vCPU · 8GB RAM
50GB disk"] m2["master2
4 vCPU · 8GB RAM
50GB disk"] m3["master3
4 vCPU · 8GB RAM
50GB disk"] end block:WK["Worker Nodes"]:1 w1["worker1
4 vCPU · 8GB RAM
50GB + 100GB raw"] w2["worker2
4 vCPU · 8GB RAM
50GB + 100GB raw"] w3["worker3
4 vCPU · 8GB RAM
50GB + 100GB raw"] end block:LB["Load Balancer"]:1 lb["lb
2 vCPU · 2GB RAM
20GB disk
HAProxy + keepalived"] end end style CP fill:#1e3a5f,stroke:#3b82f6,color:#e2e8f0 style WK fill:#1e3a5f,stroke:#10b981,color:#e2e8f0 style LB fill:#1e3a5f,stroke:#f59e0b,color:#e2e8f0

Total: ~26 vCPU, 58GB RAM, 520GB disk

Option B: Cloud VMs (AWS/GCP/Hetzner)

7 VMs tương đương cấu hình bên trên
Estimated cost: ~$200-400/tháng (Hetzner rẻ nhất)
Khuyến nghị: Hetzner Dedicated hoặc Proxmox VE

Option C: Bare-metal (Production-like)

3× Dell PowerEdge R640 hoặc tương đương:
  - 2× 16-core Xeon, 128GB RAM, 2× 480GB SSD (OS) + 4× 2TB NVMe (Ceph)
  - 4× 25GbE NICs (bonding)

4.2. Network Layout for Lab


graph TB
    subgraph MGMT["🌐 Management Network — 192.168.1.0/24"]
        direction LR
        lb["lb
192.168.1.10"] m1["master1
192.168.1.11"] m2["master2
192.168.1.12"] m3["master3
192.168.1.13"] w1["worker1
192.168.1.21"] w2["worker2
192.168.1.22"] w3["worker3
192.168.1.23"] VIP["🔷 VIP
192.168.1.100
K8s API Server"] end subgraph INTERNAL["🔒 Internal Networks"] POD["Pod Network
10.244.0.0/16
Cilium CNI"] SVC["Service Network
10.96.0.0/12
ClusterIP"] LB_POOL["MetalLB Pool
192.168.1.200–250
External Services"] end VIP --> m1 & m2 & m3 lb --> VIP style VIP fill:#dc2626,stroke:#fca5a5,color:#fff style POD fill:#1e3a5f,stroke:#3b82f6,color:#e2e8f0 style SVC fill:#1e3a5f,stroke:#10b981,color:#e2e8f0 style LB_POOL fill:#1e3a5f,stroke:#f59e0b,color:#e2e8f0

4.3. Create VMs Fast with Vagrant (Optional)

# Vagrantfile
Vagrant.configure("2") do |config|
  config.vm.box = "ubuntu/noble64"  # Ubuntu 24.04

Load Balancer

config.vm.define "lb" do |lb| lb.vm.hostname = "lb" lb.vm.network "private_network", ip: "192.168.1.10" lb.vm.provider "virtualbox" do |v| v.memory = 2048 v.cpus = 2 end end

Control Plane nodes

(1..3).each do |i| config.vm.define "master#{i}" do |master| master.vm.hostname = "master#{i}" master.vm.network "private_network", ip: "192.168.1.#{10 + i}" master.vm.provider "virtualbox" do |v| v.memory = 8192 v.cpus = 4 end end end

Worker nodes

(1..3).each do |i| config.vm.define "worker#{i}" do |worker| worker.vm.hostname = "worker#{i}" worker.vm.network "private_network", ip: "192.168.1.#{20 + i}" worker.vm.provider "virtualbox" do |v| v.memory = 8192 v.cpus = 4 # Raw disk cho Ceph OSD unless File.exist?("ceph-osd-worker#{i}.vdi") v.customize ['createmedium', 'disk', '--filename', "ceph-osd-worker#{i}.vdi", '--size', 102400] end v.customize ['storageattach', :id, '--storagectl', 'SCSI', '--port', 2, '--type', 'hdd', '--medium', "ceph-osd-worker#{i}.vdi"] end end end end

# Khởi tạo toàn bộ lab
vagrant up

# SSH vào master1
vagrant ssh master1

# Kiểm tra connectivity
for i in 10 11 12 13 21 22 23; do
  ping -c 1 192.168.1.$i
done

4.4. Configure SSH Keys for all nodes

# Trên máy workstation/jump host
ssh-keygen -t ed25519 -C "k8s-lab-admin" -f ~/.ssh/k8s-lab

Copy public key sang tất cả nodes

for host in lb master{1..3} worker{1..3}; do ssh-copy-id -i ~/.ssh/k8s-lab.pub user@${host} done

Tạo SSH config cho tiện

cat >> ~/.ssh/config << 'EOF' Host lb HostName 192.168.1.10 User root

Host master1 HostName 192.168.1.11 User root

Host master2 HostName 192.168.1.12 User root

Host master3 HostName 192.168.1.13 User root

Host worker1 HostName 192.168.1.21 User root

Host worker2 HostName 192.168.1.22 User root

Host worker3 HostName 192.168.1.23 User root

Host master* worker* lb IdentityFile ~/.ssh/k8s-lab StrictHostKeyChecking no EOF


PART 5: LEARNING ROUTE 50 LESSON

5.1. Dependency Graph between sections


graph TD
    P1["📐 Phase 1: Foundation
Bài 1-4
Chuẩn bị hạ tầng cơ bản"] P2["☸ Phase 2: K8s HA
Bài 5-10
Dựng Kubernetes HA cluster"] P3["💿 Phase 3: Rook-Ceph
Bài 11-15
Distributed Storage"] P4["🐘 Phase 4: PostgreSQL HA
Bài 16-20"] P5["📨 Phase 5: MQ HA
Bài 21-23
RabbitMQ · Kafka · Redis"] P6["🔗 Phase 6: Istio
Bài 24-27
Service Mesh"] P7["🔄 Phase 7: GitOps
Bài 28-31
ArgoCD + Helm + Vault"] P8["📊 Phase 8: Observability
Bài 32-35"] P9["🛡️ Phase 9: Security
Bài 36-39"] P10["🚀 Phase 10: Deployment Patterns
Bài 40-43"] P11["💥 Phase 11: DR & Chaos
Bài 44-45"] P12["🏭 Phase 12: Operations + Capstone
Bài 46-50"] P1 --> P2 P2 --> P3 & P5 & P6 P3 --> P4 P4 & P5 & P6 --> P7 P7 --> P8 & P9 P8 & P9 --> P10 P10 --> P11 P11 --> P12 style P1 fill:#1e40af,stroke:#3b82f6,color:#fff style P2 fill:#1e40af,stroke:#3b82f6,color:#fff style P12 fill:#15803d,stroke:#22c55e,color:#fff

5.2. Estimated time

Section Post number Time Timeline (2 hours/day)
Part 1: Foundation 4 ~8h Week 1
Part 2: K8s HA 6 ~14h Weeks 2-3
Part 3: Rook-Ceph 5 ~11h Week 3-4
Part 4: PostgreSQL__HTMLTAG_847___ 5 ~12h Week 5-6
Part 5: MQ HA 3 ~8h Week 6-7
Part 6: Istio 4 ~10h Week 7-8
Part 7: GitOps 4 ~11h Week 9-10
Part 8: Observability__HTMLTAG_887___ 4 ~10h Week 10-11
Part 9: Security 4 ~10h Week 12-13
Part 10: Deployment 4 ~9h Week 13-14
Part 11: DR 2 ~5h Week 15
Part 12: Operations 5 ~15h Week 15-18
TOTAL 50 ~123h ~18 weeks

PART 6: CONVENTIONS AND CONVENTIONS IN THE COURSE

6.1. Naming Conventions

# Namespace naming
production:     prod-<service-name>     # prod-user-service
staging:        stg-<service-name>
infrastructure: infra-<component>       # infra-monitoring, infra-storage
platform:       platform-<component>    # platform-argocd, platform-vault

Helm release naming

<component>-<environment> # postgresql-prod, redis-stg

Label standards

app.kubernetes.io/name: <service-name> app.kubernetes.io/version: <version> app.kubernetes.io/component: <component> app.kubernetes.io/part-of: <system-name> app.kubernetes.io/managed-by: helm

6.2. Symbols in the lesson

  • 💡 Tip: Useful tips, best practice
  • ⚠️ Warning: Be careful, it may cause errors
  • ❌ Danger: Absolutely do not do it in production
  • 📋 Checklist: Checklist
  • 🔬 Deep Dive: Technical explanation
  • 🛠️ Lab: Practice

💡 KEY TAKEAWAYS

  1. On-premises microservices suitable for organizations that need data sovereignty, predictable cost, and ultra-low latency
  2. Kubernetes HA is an orchestration platform, combined with the CNCF tools ecosystem to create a production platform
  3. Full stack__HTMLTAG_1003___ includes 6 layers: Infrastructure → Storage → Data → Networking → Platform → Security
  4. Each technology is selectedbased on criteria: production-grade, CNCF backed, community active
  5. Lab environment requires a minimum of 7 VMs (3 masters + 3 workers + 1 LB) with ~58GB total RAM__HTMLTAG_1012___

🎯 EXERCISE

Exercise 1: Assessing infrastructure requirements__HTMLTAG_1018___

For scenario: Fintech company needs to deploy 20 microservices, handle 10,000 requests/second, store 500GB of data, require PCI-DSS compliance.

  • Calculate the number of nodes needed (control plane, workers, storage)
  • Estimated total CPU, RAM, Storage
  • Draw network topology diagram
  • List the required components from the above stack

Exercise 2: Setup Lab Environment__HTMLTAG_1032___
  • Create 7 VMs using Option A or Option B
  • Configuring networking between VMs
  • Setup SSH key-based authentication
  • Verify ping connectivity between all nodes
  • Record the IP and hostname of each VM

Exercise 3: Comparing technology__HTMLTAG_1046___

Research and compare in detail two pairs of technologies:

  • Cilium vs Calico: Performance benchmarks, features, community
  • Rook-Ceph vs Longhorn: Scalability, features, operational complexity

📚 NEXT POST

In Lesson 2: Hardware Planning and Network Topology, we will dive into detailed sizing calculations for CPU/RAM/Disk, network topology design with VLAN, bonding, and MTU for production environment.