Chuyển đến nội dung chính

第 12 課:服務發現與註冊

用戶端與伺服器端發現、服務註冊表(Consul、etcd)、Kubernetes 基於 DNS 的發現、運行狀況檢查、負載平衡演算法和服務端點管理。

🏗️ 建築 — 第 12 課 第 12 課:服務發現與註冊

雲端原生微服務架構

第 4 部分:服務網格和網絡

亞洲開發網

第 12 課:服務發現與註冊

簡介

在微服務中,服務實例可以動態伸縮,IP不斷變化。無法對位址進行硬編碼。 服務發現是一種允許服務自動發現彼此的機制。


1. 服務發現模式

1.1 客戶端發現

客戶端查詢ServiceRegistry取得實例列表,選擇呼叫哪個實例:

┌─────────┐    ┌──────────────────┐
│ Client  │───▶│ Service Registry │
│ Service │    │ (Consul/etcd)    │
│         │◀───│                  │
│         │    │ Returns:         │
│         │    │ - 10.0.1.5:8080  │
│         │    │ - 10.0.1.6:8080  │
│         │    │ - 10.0.1.7:8080  │
│         │    └──────────────────┘
│         │
│  Client-side │
│  Load Balancer│  ← Round Robin / Random / Least Connections
│         │
│         │───▶ 10.0.1.6:8080 (chosen instance)
└─────────┘

優點:沒有代理瓶頸,客戶端自己決定路由 缺點:發現邏輯位於每個服務中(每種語言需要自己的函式庫)

1.2 伺服器端發現

用戶端發送請求到負載平衡器/路由器,路由器查詢註冊表並轉發:

┌─────────┐    ┌──────────────┐    ┌──────────────────┐
│ Client  │───▶│ Load Balancer│───▶│ Service Registry │
│ Service │    │ / Router     │◀───│                  │
│         │◀───│              │    └──────────────────┘
│         │    │              │
│         │    │              │───▶ 10.0.1.5:8080
└─────────┘    └──────────────┘

優點:簡單、與語言無關的客戶端 缺點:負載平衡器是潛在的瓶頸和單點故障

1.3 比較

標準客戶端伺服器端
客戶複雜性高(需要庫)低
跳數1(直接)2(透過LB)
語言支援需要每種語言的函式庫與語言無關
負載平衡器不需要需要(潛在的 SPOF)
範例Netflix 尤里卡 + 功能區Kubernetes 服務、AWS ELB

2.服務註冊中心

2.1 領事

HashiCorp Consul 提供服務發現、健康檢查和 KV 儲存:

┌────────────────────────────────────────────┐
│              Consul Cluster                 │
│                                             │
│  ┌──────────┐ ┌──────────┐ ┌──────────┐   │
│  │  Server  │ │  Server  │ │  Server  │   │
│  │  (Leader)│ │(Follower)│ │(Follower)│   │
│  └──────────┘ └──────────┘ └──────────┘   │
│        Raft Consensus Protocol              │
│                                             │
│  Service Catalog:                           │
│  ├── order-service                          │
│  │   ├── 10.0.1.5:8080 (passing)           │
│  │   ├── 10.0.1.6:8080 (passing)           │
│  │   └── 10.0.1.7:8080 (critical)          │
│  ├── payment-service                        │
│  │   ├── 10.0.2.3:8080 (passing)           │
│  │   └── 10.0.2.4:8080 (passing)           │
│  └── inventory-service                      │
│      └── 10.0.3.1:8080 (passing)           │
└────────────────────────────────────────────┘

服務註冊:

{
  "service": {
    "name": "order-service",
    "id": "order-service-1",
    "port": 8080,
    "tags": ["v1", "production"],
    "meta": {
      "version": "1.2.0",
      "protocol": "http"
    },
    "check": {
      "http": "http://localhost:8080/health",
      "interval": "10s",
      "timeout": "3s",
      "deregister_critical_service_after": "30s"
    }
  }
}

服務發現查詢:

# DNS interface
dig @127.0.0.1 -p 8600 order-service.service.consul SRV

# HTTP API
curl http://consul:8500/v1/health/service/order-service?passing=true

# Response
[
  {
    "Service": {
      "ID": "order-service-1",
      "Address": "10.0.1.5",
      "Port": 8080,
      "Tags": ["v1", "production"]
    },
    "Checks": [{ "Status": "passing" }]
  }
]

2.2 etcd

分散式 KV 存儲,由 Kubernetes 用於叢集狀態:

# Register service
etcdctl put /services/order-service/instances/1 \
  '{"host":"10.0.1.5","port":8080,"status":"healthy"}'

# Discover service (prefix query)
etcdctl get /services/order-service/instances/ --prefix

# Watch for changes
etcdctl watch /services/order-service/instances/ --prefix

3. Kubernetes 基於 DNS 的發現

3.1 核心DNS

Kubernetes 透過 CoreDNS 內建服務發現:

┌───────────────────────────────────────────────┐
│            Kubernetes Cluster                  │
│                                                │
│  ┌──────────────┐                              │
│  │   CoreDNS    │ ← Watches Kubernetes API     │
│  │   (kube-dns) │                              │
│  └──────┬───────┘                              │
│         │                                      │
│  DNS Records:                                  │
│  ├── order-service.default.svc.cluster.local   │
│  │   → ClusterIP: 10.96.45.32                 │
│  ├── payment-service.default.svc.cluster.local │
│  │   → ClusterIP: 10.96.78.91                 │
│  └── order-service.staging.svc.cluster.local   │
│      → ClusterIP: 10.96.12.55                 │
│                                                │
│  Format: <service>.<namespace>.svc.cluster.local│
└───────────────────────────────────────────────┘

3.2 Kubernetes 服務類型

# ClusterIP (default) — internal only
apiVersion: v1
kind: Service
metadata:
  name: order-service
  namespace: default
spec:
  type: ClusterIP
  selector:
    app: order-service
  ports:
    - port: 8080
      targetPort: 8080

---
# Headless Service — returns Pod IPs directly (no load balancing)
apiVersion: v1
kind: Service
metadata:
  name: order-service-headless
spec:
  clusterIP: None  # ← Headless
  selector:
    app: order-service
  ports:
    - port: 8080
# ClusterIP service → resolves to virtual IP
nslookup order-service.default.svc.cluster.local
# → 10.96.45.32

# Headless service → resolves to all Pod IPs
nslookup order-service-headless.default.svc.cluster.local
# → 10.0.1.5, 10.0.1.6, 10.0.1.7

# Within same namespace, short name works
curl http://order-service:8080/api/orders

# Cross-namespace
curl http://order-service.staging:8080/api/orders

3.3 端點切片

# Kubernetes tự động tạo EndpointSlice cho mỗi Service
apiVersion: discovery.k8s.io/v1
kind: EndpointSlice
metadata:
  name: order-service-abc12
  labels:
    kubernetes.io/service-name: order-service
addressType: IPv4
endpoints:
  - addresses: ["10.0.1.5"]
    conditions:
      ready: true
      serving: true
  - addresses: ["10.0.1.6"]
    conditions:
      ready: true
      serving: true
  - addresses: ["10.0.1.7"]
    conditions:
      ready: false   # Not ready — excluded from routing
      serving: false
ports:
  - port: 8080
    protocol: TCP

4. 健康檢查

4.1 健康檢查類型

┌─────────────────────────────────────────────────┐
│              Health Check Levels                 │
│                                                  │
│  ┌─────────────────────────────────────────┐    │
│  │ Liveness: "Is the process alive?"       │    │
│  │ → Fail: Restart container               │    │
│  │ → Check: process not deadlocked         │    │
│  └─────────────────────────────────────────┘    │
│                                                  │
│  ┌─────────────────────────────────────────┐    │
│  │ Readiness: "Can it handle requests?"    │    │
│  │ → Fail: Remove from Service endpoints   │    │
│  │ → Check: DB connected, cache warm       │    │
│  └─────────────────────────────────────────┘    │
│                                                  │
│  ┌─────────────────────────────────────────┐    │
│  │ Startup: "Has it finished starting?"    │    │
│  │ → Fail: Keep waiting (don't kill early) │    │
│  │ → Check: initialization complete        │    │
│  └─────────────────────────────────────────┘    │
└─────────────────────────────────────────────────┘

4.2 實施

@RestController
public class HealthController {

    @Autowired
    private DataSource dataSource;
    
    @Autowired
    private RedisTemplate<String, String> redis;

    // Liveness — lightweight, no dependency check
    @GetMapping("/health/live")
    public ResponseEntity<Map<String, String>> liveness() {
        return ResponseEntity.ok(Map.of("status", "UP"));
    }

    // Readiness — check dependencies
    @GetMapping("/health/ready")
    public ResponseEntity<Map<String, Object>> readiness() {
        Map<String, Object> checks = new HashMap<>();
        boolean ready = true;

        // Check database
        try {
            dataSource.getConnection().isValid(2);
            checks.put("database", "UP");
        } catch (Exception e) {
            checks.put("database", "DOWN");
            ready = false;
        }

        // Check Redis
        try {
            redis.getConnectionFactory().getConnection().ping();
            checks.put("redis", "UP");
        } catch (Exception e) {
            checks.put("redis", "DOWN");
            ready = false;
        }

        checks.put("status", ready ? "UP" : "DOWN");
        return ready 
            ? ResponseEntity.ok(checks)
            : ResponseEntity.status(503).body(checks);
    }
}

5.負載平衡演算法

5.1 流行演算法

Round Robin:
  Request 1 → Instance A
  Request 2 → Instance B
  Request 3 → Instance C
  Request 4 → Instance A (vòng lại)

Weighted Round Robin:
  Instance A (weight: 3) → nhận 3/6 requests
  Instance B (weight: 2) → nhận 2/6 requests
  Instance C (weight: 1) → nhận 1/6 requests

Least Connections:
  Chọn instance có ít connection nhất hiện tại
  Phù hợp khi request có duration khác nhau

Random:
  Chọn ngẫu nhiên — đơn giản, hiệu quả cho large pool

Consistent Hashing:
  Hash request key → map đến instance cố định
  Phù hợp cho sticky sessions, caching scenarios

Power of Two Choices:
  Random chọn 2 instances, pick instance ít load hơn
  Tốt hơn pure random, ít overhead hơn least connections

5.2 Kubernetes kube-proxy 模式

iptables mode (default):
  → Random selection via iptables rules
  → No health-aware routing

IPVS mode:
  → Supports multiple algorithms
  → rr (Round Robin), lc (Least Connection), 
    dh (Destination Hashing), sh (Source Hashing)
  → Better performance at scale

# Enable IPVS mode
kubectl edit configmap kube-proxy -n kube-system
# mode: "ipvs"
# ipvs:
#   scheduler: "lc"  # Least Connections

6. 最佳實踐

1. Sử dụng Kubernetes DNS cho intra-cluster discovery
   → Không cần external registry

2. Headless Service cho stateful workloads
   → Client-side load balancing cho gRPC

3. Health checks luôn bao gồm cả liveness + readiness
   → Liveness: lightweight, không check dependencies
   → Readiness: check downstream dependencies

4. Graceful shutdown với preStop hook
   → Deregister trước khi shutdown
   → Drain existing connections

5. Service mesh thay thế client-side discovery
   → Sidecar proxy xử lý routing + load balancing
   → Application code không cần discovery logic
# Graceful shutdown example
spec:
  terminationGracePeriodSeconds: 30
  containers:
    - name: app
      lifecycle:
        preStop:
          exec:
            command: ["/bin/sh", "-c", "sleep 5"]  # Wait for endpoint removal

總結

  • 客戶端發現:客戶端查詢註冊表,自我負載平衡
  • 伺服器端發現:路由器/LB查詢註冊表,轉發請求
  • Kubernetes DNS:內建發現,最適合 K8s 工作負載
  • 健康檢查:Liveness(還活著?)、Readiness(準備好?)、Startup(開始了嗎?)
  • 在 Kubernetes 中,Kubernetes Service + CoreDNS 通常就足夠了,不需要單獨的 Consul/etcd
  • Service Mesh (Istio/Linkerd) 透過智慧路由升級發現