
簡介
在微服務中,服務實例可以動態伸縮,IP不斷變化。無法對位址進行硬編碼。 服務發現是一種允許服務自動發現彼此的機制。
1. 服務發現模式
1.1 客戶端發現
客戶端查詢ServiceRegistry取得實例列表,選擇呼叫哪個實例:
┌─────────┐ ┌──────────────────┐
│ Client │───▶│ Service Registry │
│ Service │ │ (Consul/etcd) │
│ │◀───│ │
│ │ │ Returns: │
│ │ │ - 10.0.1.5:8080 │
│ │ │ - 10.0.1.6:8080 │
│ │ │ - 10.0.1.7:8080 │
│ │ └──────────────────┘
│ │
│ Client-side │
│ Load Balancer│ ← Round Robin / Random / Least Connections
│ │
│ │───▶ 10.0.1.6:8080 (chosen instance)
└─────────┘
優點:沒有代理瓶頸,客戶端自己決定路由 缺點:發現邏輯位於每個服務中(每種語言需要自己的函式庫)
1.2 伺服器端發現
用戶端發送請求到負載平衡器/路由器,路由器查詢註冊表並轉發:
┌─────────┐ ┌──────────────┐ ┌──────────────────┐
│ Client │───▶│ Load Balancer│───▶│ Service Registry │
│ Service │ │ / Router │◀───│ │
│ │◀───│ │ └──────────────────┘
│ │ │ │
│ │ │ │───▶ 10.0.1.5:8080
└─────────┘ └──────────────┘
優點:簡單、與語言無關的客戶端 缺點:負載平衡器是潛在的瓶頸和單點故障
1.3 比較
| 標準 | 客戶端 | 伺服器端 |
|---|---|---|
| 客戶複雜性 | 高(需要庫) | 低 |
| 跳數 | 1(直接) | 2(透過LB) |
| 語言支援 | 需要每種語言的函式庫 | 與語言無關 |
| 負載平衡器 | 不需要 | 需要(潛在的 SPOF) |
| 範例 | Netflix 尤里卡 + 功能區 | Kubernetes 服務、AWS ELB |
2.服務註冊中心
2.1 領事
HashiCorp Consul 提供服務發現、健康檢查和 KV 儲存:
┌────────────────────────────────────────────┐
│ Consul Cluster │
│ │
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
│ │ Server │ │ Server │ │ Server │ │
│ │ (Leader)│ │(Follower)│ │(Follower)│ │
│ └──────────┘ └──────────┘ └──────────┘ │
│ Raft Consensus Protocol │
│ │
│ Service Catalog: │
│ ├── order-service │
│ │ ├── 10.0.1.5:8080 (passing) │
│ │ ├── 10.0.1.6:8080 (passing) │
│ │ └── 10.0.1.7:8080 (critical) │
│ ├── payment-service │
│ │ ├── 10.0.2.3:8080 (passing) │
│ │ └── 10.0.2.4:8080 (passing) │
│ └── inventory-service │
│ └── 10.0.3.1:8080 (passing) │
└────────────────────────────────────────────┘
服務註冊:
{
"service": {
"name": "order-service",
"id": "order-service-1",
"port": 8080,
"tags": ["v1", "production"],
"meta": {
"version": "1.2.0",
"protocol": "http"
},
"check": {
"http": "http://localhost:8080/health",
"interval": "10s",
"timeout": "3s",
"deregister_critical_service_after": "30s"
}
}
}
服務發現查詢:
# DNS interface
dig @127.0.0.1 -p 8600 order-service.service.consul SRV
# HTTP API
curl http://consul:8500/v1/health/service/order-service?passing=true
# Response
[
{
"Service": {
"ID": "order-service-1",
"Address": "10.0.1.5",
"Port": 8080,
"Tags": ["v1", "production"]
},
"Checks": [{ "Status": "passing" }]
}
]
2.2 etcd
分散式 KV 存儲,由 Kubernetes 用於叢集狀態:
# Register service
etcdctl put /services/order-service/instances/1 \
'{"host":"10.0.1.5","port":8080,"status":"healthy"}'
# Discover service (prefix query)
etcdctl get /services/order-service/instances/ --prefix
# Watch for changes
etcdctl watch /services/order-service/instances/ --prefix
3. Kubernetes 基於 DNS 的發現
3.1 核心DNS
Kubernetes 透過 CoreDNS 內建服務發現:
┌───────────────────────────────────────────────┐
│ Kubernetes Cluster │
│ │
│ ┌──────────────┐ │
│ │ CoreDNS │ ← Watches Kubernetes API │
│ │ (kube-dns) │ │
│ └──────┬───────┘ │
│ │ │
│ DNS Records: │
│ ├── order-service.default.svc.cluster.local │
│ │ → ClusterIP: 10.96.45.32 │
│ ├── payment-service.default.svc.cluster.local │
│ │ → ClusterIP: 10.96.78.91 │
│ └── order-service.staging.svc.cluster.local │
│ → ClusterIP: 10.96.12.55 │
│ │
│ Format: <service>.<namespace>.svc.cluster.local│
└───────────────────────────────────────────────┘
3.2 Kubernetes 服務類型
# ClusterIP (default) — internal only
apiVersion: v1
kind: Service
metadata:
name: order-service
namespace: default
spec:
type: ClusterIP
selector:
app: order-service
ports:
- port: 8080
targetPort: 8080
---
# Headless Service — returns Pod IPs directly (no load balancing)
apiVersion: v1
kind: Service
metadata:
name: order-service-headless
spec:
clusterIP: None # ← Headless
selector:
app: order-service
ports:
- port: 8080
# ClusterIP service → resolves to virtual IP
nslookup order-service.default.svc.cluster.local
# → 10.96.45.32
# Headless service → resolves to all Pod IPs
nslookup order-service-headless.default.svc.cluster.local
# → 10.0.1.5, 10.0.1.6, 10.0.1.7
# Within same namespace, short name works
curl http://order-service:8080/api/orders
# Cross-namespace
curl http://order-service.staging:8080/api/orders
3.3 端點切片
# Kubernetes tự động tạo EndpointSlice cho mỗi Service
apiVersion: discovery.k8s.io/v1
kind: EndpointSlice
metadata:
name: order-service-abc12
labels:
kubernetes.io/service-name: order-service
addressType: IPv4
endpoints:
- addresses: ["10.0.1.5"]
conditions:
ready: true
serving: true
- addresses: ["10.0.1.6"]
conditions:
ready: true
serving: true
- addresses: ["10.0.1.7"]
conditions:
ready: false # Not ready — excluded from routing
serving: false
ports:
- port: 8080
protocol: TCP
4. 健康檢查
4.1 健康檢查類型
┌─────────────────────────────────────────────────┐
│ Health Check Levels │
│ │
│ ┌─────────────────────────────────────────┐ │
│ │ Liveness: "Is the process alive?" │ │
│ │ → Fail: Restart container │ │
│ │ → Check: process not deadlocked │ │
│ └─────────────────────────────────────────┘ │
│ │
│ ┌─────────────────────────────────────────┐ │
│ │ Readiness: "Can it handle requests?" │ │
│ │ → Fail: Remove from Service endpoints │ │
│ │ → Check: DB connected, cache warm │ │
│ └─────────────────────────────────────────┘ │
│ │
│ ┌─────────────────────────────────────────┐ │
│ │ Startup: "Has it finished starting?" │ │
│ │ → Fail: Keep waiting (don't kill early) │ │
│ │ → Check: initialization complete │ │
│ └─────────────────────────────────────────┘ │
└─────────────────────────────────────────────────┘
4.2 實施
@RestController
public class HealthController {
@Autowired
private DataSource dataSource;
@Autowired
private RedisTemplate<String, String> redis;
// Liveness — lightweight, no dependency check
@GetMapping("/health/live")
public ResponseEntity<Map<String, String>> liveness() {
return ResponseEntity.ok(Map.of("status", "UP"));
}
// Readiness — check dependencies
@GetMapping("/health/ready")
public ResponseEntity<Map<String, Object>> readiness() {
Map<String, Object> checks = new HashMap<>();
boolean ready = true;
// Check database
try {
dataSource.getConnection().isValid(2);
checks.put("database", "UP");
} catch (Exception e) {
checks.put("database", "DOWN");
ready = false;
}
// Check Redis
try {
redis.getConnectionFactory().getConnection().ping();
checks.put("redis", "UP");
} catch (Exception e) {
checks.put("redis", "DOWN");
ready = false;
}
checks.put("status", ready ? "UP" : "DOWN");
return ready
? ResponseEntity.ok(checks)
: ResponseEntity.status(503).body(checks);
}
}
5.負載平衡演算法
5.1 流行演算法
Round Robin:
Request 1 → Instance A
Request 2 → Instance B
Request 3 → Instance C
Request 4 → Instance A (vòng lại)
Weighted Round Robin:
Instance A (weight: 3) → nhận 3/6 requests
Instance B (weight: 2) → nhận 2/6 requests
Instance C (weight: 1) → nhận 1/6 requests
Least Connections:
Chọn instance có ít connection nhất hiện tại
Phù hợp khi request có duration khác nhau
Random:
Chọn ngẫu nhiên — đơn giản, hiệu quả cho large pool
Consistent Hashing:
Hash request key → map đến instance cố định
Phù hợp cho sticky sessions, caching scenarios
Power of Two Choices:
Random chọn 2 instances, pick instance ít load hơn
Tốt hơn pure random, ít overhead hơn least connections
5.2 Kubernetes kube-proxy 模式
iptables mode (default):
→ Random selection via iptables rules
→ No health-aware routing
IPVS mode:
→ Supports multiple algorithms
→ rr (Round Robin), lc (Least Connection),
dh (Destination Hashing), sh (Source Hashing)
→ Better performance at scale
# Enable IPVS mode
kubectl edit configmap kube-proxy -n kube-system
# mode: "ipvs"
# ipvs:
# scheduler: "lc" # Least Connections
6. 最佳實踐
1. Sử dụng Kubernetes DNS cho intra-cluster discovery
→ Không cần external registry
2. Headless Service cho stateful workloads
→ Client-side load balancing cho gRPC
3. Health checks luôn bao gồm cả liveness + readiness
→ Liveness: lightweight, không check dependencies
→ Readiness: check downstream dependencies
4. Graceful shutdown với preStop hook
→ Deregister trước khi shutdown
→ Drain existing connections
5. Service mesh thay thế client-side discovery
→ Sidecar proxy xử lý routing + load balancing
→ Application code không cần discovery logic
# Graceful shutdown example
spec:
terminationGracePeriodSeconds: 30
containers:
- name: app
lifecycle:
preStop:
exec:
command: ["/bin/sh", "-c", "sleep 5"] # Wait for endpoint removal
總結
- 客戶端發現:客戶端查詢註冊表,自我負載平衡
- 伺服器端發現:路由器/LB查詢註冊表,轉發請求
- Kubernetes DNS:內建發現,最適合 K8s 工作負載
- 健康檢查:Liveness(還活著?)、Readiness(準備好?)、Startup(開始了嗎?)
- 在 Kubernetes 中,Kubernetes Service + CoreDNS 通常就足夠了,不需要單獨的 Consul/etcd
- Service Mesh (Istio/Linkerd) 透過智慧路由升級發現