Chuyển đến nội dung chính

BÀI 23: REDIS HA — SENTINEL VÀ CLUSTER MODE

Deploy Redis HA trên Kubernetes với hai mode: Sentinel (master-replica) và Cluster (sharding), caching strategies, persistence, monitoring và best practices.

🔒 DevSecOps — Bài 23 BÀI 23: REDIS HA — SENTINEL VÀ CLUSTER MODE

Deploy Microservices On-Premises với Kubernetes HA

Phần 5: Message Queue HA (RabbitMQ, Kafka, Redis)

xdev.asia

🎯 MỤC TIÊU BÀI HỌC

  • ✅ Hiểu Redis Sentinel vs Cluster mode — khi nào dùng gì
  • ✅ Deploy Redis Sentinel HA trên Kubernetes
  • ✅ Deploy Redis Cluster mode (sharding)
  • ✅ Cấu hình persistence: RDB vs AOF
  • ✅ Caching strategies và best practices
  • ✅ Monitoring Redis với Prometheus

PHẦN 1: REDIS HA — SENTINEL vs CLUSTER


Redis Sentinel Mode (Master-Replica + Failover):

┌─────────────┐     ┌─────────────┐     ┌─────────────┐
│ Sentinel 1  │     │ Sentinel 2  │     │ Sentinel 3  │
│ (Quorum)    │◄───►│ (Quorum)    │◄───►│ (Quorum)    │
└──────┬──────┘     └──────┬──────┘     └──────┬──────┘
       │                   │                   │
       ▼                   ▼                   ▼
┌──────────┐        ┌──────────┐        ┌──────────┐
│  MASTER  │───────►│ REPLICA 1│        │ REPLICA 2│
│  (R/W)   │───────►│  (Read)  │        │  (Read)  │
└──────────┘        └──────────┘        └──────────┘

Redis Cluster Mode (Sharding):

┌─────────────────────────────────────────────────────┐
│                 16384 Hash Slots                     │
├─────────────┬─────────────┬─────────────────────────┤
│ Slots 0-5460│ Slots 5461- │ Slots 10923-16383      │
│             │ 10922       │                         │
│ ┌────────┐  │ ┌────────┐  │ ┌────────┐             │
│ │Master 1│  │ │Master 2│  │ │Master 3│             │
│ └───┬────┘  │ └───┬────┘  │ └───┬────┘             │
│     │       │     │       │     │                   │
│ ┌───▼────┐  │ ┌───▼────┐  │ ┌───▼────┐             │
│ │Replica │  │ │Replica │  │ │Replica │             │
│ │  1a    │  │ │  2a    │  │ │  3a    │             │
│ └────────┘  │ └────────┘  │ └────────┘             │
└─────────────┴─────────────┴─────────────────────────┘
FeatureSentinel ModeCluster Mode
Data DistributionAll data trên MasterSharded (hash slots)
Max Dataset SizeSingle node RAMSum of all masters' RAM
Read ScalingRead replicasRead replicas per shard
Write ScalingSingle master onlyMultiple masters (horizontal)
FailoverSentinel quorum voteBuilt-in (gossip protocol)
Multi-key OpsSupportedOnly same hash slot ({tag})
ComplexitySimpleMore complex
Best ForCache, sessions (< 32GB)Large datasets, high write throughput

PHẦN 2: DEPLOY REDIS SENTINEL HA

2.1. Install với Bitnami Helm Chart

# Add Bitnami repo:
helm repo add bitnami https://charts.bitnami.com/bitnami
helm repo update

Install Redis Sentinel:

helm install redis-sentinel bitnami/redis
--namespace caching
--create-namespace
--set architecture=replication
--set auth.password="$(openssl rand -base64 32)"
--set sentinel.enabled=true
--set sentinel.quorum=2
--set replica.replicaCount=3
--set sentinel.resources.requests.cpu=100m
--set sentinel.resources.requests.memory=128Mi
--set master.persistence.storageClass=ceph-block
--set master.persistence.size=10Gi
--set replica.persistence.storageClass=ceph-block
--set replica.persistence.size=10Gi
--set master.resources.requests.cpu=250m
--set master.resources.requests.memory=512Mi
--set master.resources.limits.cpu=1
--set master.resources.limits.memory=1Gi
--set metrics.enabled=true
--set metrics.serviceMonitor.enabled=true

2.2. Custom Values File (chi tiết)

# redis-sentinel-values.yaml:
architecture: replication

auth: enabled: true existingSecret: redis-secret existingSecretPasswordKey: password

master: count: 1 persistence: enabled: true storageClass: ceph-block size: 10Gi resources: requests: cpu: 250m memory: 512Mi limits: cpu: "1" memory: 1Gi configuration: | maxmemory 768mb maxmemory-policy allkeys-lru save 900 1 save 300 10 save 60 10000 appendonly yes appendfsync everysec no-appendfsync-on-rewrite yes auto-aof-rewrite-percentage 100 auto-aof-rewrite-min-size 64mb tcp-keepalive 300 timeout 0 hz 10 dynamic-hz yes

replica: replicaCount: 2 persistence: enabled: true storageClass: ceph-block size: 10Gi resources: requests: cpu: 250m memory: 512Mi limits: cpu: "1" memory: 1Gi

sentinel: enabled: true quorum: 2 downAfterMilliseconds: 5000 failoverTimeout: 60000 resources: requests: cpu: 100m memory: 128Mi

metrics: enabled: true serviceMonitor: enabled: true namespace: monitoring resources: requests: cpu: 50m memory: 64Mi

podAntiAffinityPreset: hard

# Deploy with values:
kubectl create namespace caching

kubectl -n caching create secret generic redis-secret \
  --from-literal=password="$(openssl rand -base64 32)"

helm install redis-sentinel bitnami/redis \
  --namespace caching \
  -f redis-sentinel-values.yaml

# Verify:
kubectl -n caching get pods
# redis-sentinel-node-0   3/3   Running   (master + sentinel + metrics)
# redis-sentinel-node-1   3/3   Running   (replica + sentinel + metrics)
# redis-sentinel-node-2   3/3   Running   (replica + sentinel + metrics)

# Check sentinel:
kubectl -n caching exec redis-sentinel-node-0 -c sentinel -- \
  redis-cli -p 26379 sentinel masters
# name: mymaster
# ip: redis-sentinel-node-0.redis-sentinel-headless.caching
# port: 6379
# flags: master
# num-slaves: 2
# num-other-sentinels: 2
# quorum: 2

PHẦN 3: DEPLOY REDIS CLUSTER MODE

# redis-cluster-values.yaml:
# Cho workloads cần horizontal write scaling
architecture: cluster

cluster:
  nodes: 6               # 3 masters + 3 replicas
  replicas: 1             # 1 replica per master

auth:
  enabled: true
  existingSecret: redis-cluster-secret

persistence:
  enabled: true
  storageClass: ceph-block
  size: 10Gi

resources:
  requests:
    cpu: 250m
    memory: 512Mi
  limits:
    cpu: "1"
    memory: 1Gi

redis:
  configmap: |
    maxmemory 768mb
    maxmemory-policy allkeys-lru
    cluster-require-full-coverage no
    cluster-allow-reads-when-down yes

metrics:
  enabled: true
  serviceMonitor:
    enabled: true

podAntiAffinityPreset: hard
# Deploy Redis Cluster:
helm install redis-cluster bitnami/redis-cluster \
  --namespace caching \
  -f redis-cluster-values.yaml

# Verify cluster:
kubectl -n caching exec redis-cluster-0 -- \
  redis-cli -a "$REDIS_PASSWORD" cluster info
# cluster_state:ok
# cluster_slots_assigned:16384
# cluster_slots_ok:16384
# cluster_size:3
# cluster_known_nodes:6

# Check slot distribution:
kubectl -n caching exec redis-cluster-0 -- \
  redis-cli -a "$REDIS_PASSWORD" cluster nodes
# node-id master 10.244.1.x:6379 0-5460
# node-id master 10.244.2.x:6379 5461-10922
# node-id master 10.244.3.x:6379 10923-16383

PHẦN 4: PERSISTENCE — RDB vs AOF

FeatureRDB (Snapshotting)AOF (Append-Only File)RDB + AOF
MechanismPoint-in-time snapshotLog every write operationBoth
Data LossUp to last snapshot~1 second (appendfsync everysec)Minimal
Recovery SpeedFast (load binary)Slower (replay operations)Uses RDB first
Disk I/OPeriodic burst (fork)Continuous (small writes)Both
File SizeCompactLarger (rewrite helps)Both files
Best ForCaching (acceptable loss)Session store (minimal loss)Production (recommended)
# Configuration cho production (RDB + AOF):
# save 900 1       → Snapshot nếu 1 key change trong 900s
# save 300 10      → Snapshot nếu 10 keys change trong 300s  
# save 60 10000    → Snapshot nếu 10000 keys change trong 60s
# appendonly yes
# appendfsync everysec

PHẦN 5: CACHING STRATEGIES


1. Cache-Aside (Lazy Loading):
   App ──► Cache HIT? ──► Return data
                │ MISS
                ▼
           Read from DB → Store in Cache → Return

2. Write-Through:
   App ──► Write to Cache ──► Write to DB (sync)

3. Write-Behind (Write-Back):
   App ──► Write to Cache ──► Async write to DB

4. Read-Through:
   App ──► Cache (auto-loads from DB on miss)
# Python example - Cache-Aside pattern:
import redis
import json

r = redis.Redis(
    host='redis-sentinel.caching.svc',
    port=26379,
    password=os.environ['REDIS_PASSWORD'],
    sentinel_manager=True,
    db=0,
    decode_responses=True
)

def get_user(user_id):
    cache_key = f"user:{user_id}"
    
    # Try cache first:
    cached = r.get(cache_key)
    if cached:
        return json.loads(cached)
    
    # Cache miss → query DB:
    user = db.query("SELECT * FROM users WHERE id = %s", user_id)
    
    # Store in cache with TTL:
    r.setex(cache_key, 3600, json.dumps(user))  # 1 hour TTL
    
    return user

def update_user(user_id, data):
    # Update DB first:
    db.execute("UPDATE users SET ... WHERE id = %s", user_id)
    
    # Invalidate cache:
    r.delete(f"user:{user_id}")

PHẦN 6: APPLICATION CONNECTION

# App connecting to Redis Sentinel:
apiVersion: apps/v1
kind: Deployment
metadata:
  name: my-app
spec:
  template:
    spec:
      containers:
        - name: app
          env:
            - name: REDIS_SENTINEL_HOST
              value: "redis-sentinel.caching.svc"
            - name: REDIS_SENTINEL_PORT
              value: "26379"
            - name: REDIS_MASTER_NAME
              value: "mymaster"
            - name: REDIS_PASSWORD
              valueFrom:
                secretKeyRef:
                  name: redis-app-secret
                  key: password

PHẦN 7: MONITORING

# Key Redis metrics:
# redis_connected_clients          — Current connections
# redis_used_memory_bytes          — Memory usage
# redis_evicted_keys_total         — Keys evicted (maxmemory)
# redis_keyspace_hits_total        — Cache hits
# redis_keyspace_misses_total      — Cache misses
# redis_commands_processed_total   — Commands/sec

# Cache Hit Rate formula:
# hit_rate = hits / (hits + misses) × 100%
# Target: > 90%

# Grafana dashboard: ID 11835 (Redis Dashboard for Prometheus)
# Alerting:
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
  name: redis-alerts
  namespace: caching
spec:
  groups:
    - name: redis
      rules:
        - alert: RedisHighMemory
          expr: redis_used_memory_bytes / redis_maxmemory > 0.9
          for: 5m
          labels:
            severity: warning

        - alert: RedisMasterDown
          expr: redis_up{role="master"} == 0
          for: 1m
          labels:
            severity: critical

        - alert: RedisLowCacheHitRate
          expr: |
            redis_keyspace_hits_total / 
            (redis_keyspace_hits_total + redis_keyspace_misses_total) < 0.8
          for: 15m
          labels:
            severity: warning

💡 KEY TAKEAWAYS

  1. Sentinel: Simple HA, single master, good for < 32GB datasets
  2. Cluster: Horizontal scaling, multi-master, for large datasets
  3. Persistence: RDB + AOF cho production, balance durability vs performance
  4. maxmemory-policy: allkeys-lru phổ biến nhất cho caching
  5. Cache-Aside: Pattern nên dùng, invalidate on write
  6. Monitor: Cache hit rate > 90%, memory usage, evictions

🎯 BÀI TẬP

Bài tập 1: Redis Sentinel Failover

  • Deploy Redis Sentinel 3-node
  • Write data to master, kill master pod
  • Verify Sentinel promotes replica, data intact

Bài tập 2: Benchmark

  • Run redis-benchmark: SET/GET 100,000 keys
  • Compare latency: with/without persistence
  • Monitor memory fragmentation ratio

📚 BÀI TIẾP THEO

Trong Bài 24: Kiến trúc Istio Service Mesh, chúng ta sẽ tìm hiểu service mesh và deploy Istio cho microservices communication.