Chuyển đến nội dung chính

LESSON 23: REDIS HA — SENTINEL AND CLUSTER MODE

Deploy Redis HA on Kubernetes with two modes: Sentinel (master-replica) and Cluster (sharding), caching strategies, persistence, monitoring and best practices.

🔒 DevSecOps — Lesson 23 LESSON 23: REDIS HA — SENTINEL AND CLUSTER MODE

Deploy Microservices On-Premises with Kubernetes HA

Part 5: Message Queue HA (RabbitMQ, Kafka, Redis)

xdev.asia

🎯 LESSON OBJECTIVE__HTMLTAG_68___
  • ✅ Understanding Redis Sentinel vs Cluster mode — when to use what
  • ✅ Deploy Redis Sentinel HA on Kubernetes
  • ✅ Deploy Redis Cluster mode (sharding)
  • ✅ Persistence configuration: RDB vs AOF
  • ✅ Caching strategies and best practices
  • ✅ Monitoring Redis with Prometheus

PART 1: REDIS HA — SENTINEL vs CLUSTER


Redis Sentinel Mode (Master-Replica + Failover):

┌─────────────┐     ┌─────────────┐     ┌─────────────┐
│ Sentinel 1  │     │ Sentinel 2  │     │ Sentinel 3  │
│ (Quorum)    │◄───►│ (Quorum)    │◄───►│ (Quorum)    │
└──────┬──────┘     └──────┬──────┘     └──────┬──────┘
       │                   │                   │
       ▼                   ▼                   ▼
┌──────────┐        ┌──────────┐        ┌──────────┐
│  MASTER  │───────►│ REPLICA 1│        │ REPLICA 2│
│  (R/W)   │───────►│  (Read)  │        │  (Read)  │
└──────────┘        └──────────┘        └──────────┘

Redis Cluster Mode (Sharding):

┌─────────────────────────────────────────────────────┐
│                 16384 Hash Slots                     │
├─────────────┬─────────────┬─────────────────────────┤
│ Slots 0-5460│ Slots 5461- │ Slots 10923-16383      │
│             │ 10922       │                         │
│ ┌────────┐  │ ┌────────┐  │ ┌────────┐             │
│ │Master 1│  │ │Master 2│  │ │Master 3│             │
│ └───┬────┘  │ └───┬────┘  │ └───┬────┘             │
│     │       │     │       │     │                   │
│ ┌───▼────┐  │ ┌───▼────┐  │ ┌───▼────┐             │
│ │Replica │  │ │Replica │  │ │Replica │             │
│ │  1a    │  │ │  2a    │  │ │  3a    │             │
│ └────────┘  │ └────────┘  │ └────────┘             │
└─────────────┴─────────────┴─────────────────────────┘
FeatureSentinel ModeCluster Mode
Data DistributionAll data on MasterSharded (hash slots)
Max Dataset SizeSingle node RAM_Sum of all masters' RAM_
Read Scaling_Read replicas_Read replicas per shard_
Write Scaling_Single master onlyMultiple masters (horizontal)
FailoverSentinel quorum voteBuilt-in (gossip protocol)
Multi-key OpsSupportedOnly same hash slot ({tag})
ComplexitySimpleMore complexity
Best ForCache, sessions (< 32GB)_Large datasets, high write throughput__HTMLTAG_161___

PART 2: DEPLOY REDIS SENTINEL HA

2.1. Install with Bitnami Helm Chart

# Add Bitnami repo:
helm repo add bitnami https://charts.bitnami.com/bitnami
helm repo update

Install Redis Sentinel:

helm install redis-sentinel bitnami/redis
--namespace caching
--create-namespace
--set architecture=replication
--set auth.password="$(openssl rand -base64 32)"
--set sentinel.enabled=true
--set sentinel.quorum=2
--set replica.replicaCount=3
--set sentinel.resources.requests.cpu=100m
--set sentinel.resources.requests.memory=128Mi
--set master.persistence.storageClass=ceph-block
--set master.persistence.size=10Gi
--set replica.persistence.storageClass=ceph-block
--set replica.persistence.size=10Gi
--set master.resources.requests.cpu=250m
--set master.resources.requests.memory=512Mi
--set master.resources.limits.cpu=1
--set master.resources.limits.memory=1Gi
--set metrics.enabled=true
--set metrics.serviceMonitor.enabled=true

2.2. Custom Values File (details)

# redis-sentinel-values.yaml:
architecture: replication

auth: enabled: true existingSecret: redis-secret existingSecretPasswordKey: password

master: count: 1 persistence: enabled: true storageClass: ceph-block size: 10Gi resources: requests: cpu: 250m memory: 512Mi limits: cpu: "1" memory: 1Gi configuration: | maxmemory 768mb maxmemory-policy allkeys-lru save 900 1 save 300 10 save 60 10000 appendonly yes appendfsync everysec no-appendfsync-on-rewrite yes auto-aof-rewrite-percentage 100 auto-aof-rewrite-min-size 64mb tcp-keepalive 300 timeout 0 hz 10 dynamic-hz yes

replica: replicaCount: 2 persistence: enabled: true storageClass: ceph-block size: 10Gi resources: requests: cpu: 250m memory: 512Mi limits: cpu: "1" memory: 1Gi

sentinel: enabled: true quorum: 2 downAfterMilliseconds: 5000 failoverTimeout: 60000 resources: requests: cpu: 100m memory: 128Mi

metrics: enabled: true serviceMonitor: enabled: true namespace: monitoring resources: requests: cpu: 50m memory: 64Mi

podAntiAffinityPreset: hard

# Deploy with values:
kubectl create namespace caching

kubectl -n caching create secret generic redis-secret \
  --from-literal=password="$(openssl rand -base64 32)"

helm install redis-sentinel bitnami/redis \
  --namespace caching \
  -f redis-sentinel-values.yaml

# Verify:
kubectl -n caching get pods
# redis-sentinel-node-0   3/3   Running   (master + sentinel + metrics)
# redis-sentinel-node-1   3/3   Running   (replica + sentinel + metrics)
# redis-sentinel-node-2   3/3   Running   (replica + sentinel + metrics)

# Check sentinel:
kubectl -n caching exec redis-sentinel-node-0 -c sentinel -- \
  redis-cli -p 26379 sentinel masters
# name: mymaster
# ip: redis-sentinel-node-0.redis-sentinel-headless.caching
# port: 6379
# flags: master
# num-slaves: 2
# num-other-sentinels: 2
# quorum: 2

PART 3: DEPLOY REDIS CLUSTER MODE

# redis-cluster-values.yaml:
# Cho workloads cần horizontal write scaling
architecture: cluster

cluster:
  nodes: 6               # 3 masters + 3 replicas
  replicas: 1             # 1 replica per master

auth:
  enabled: true
  existingSecret: redis-cluster-secret

persistence:
  enabled: true
  storageClass: ceph-block
  size: 10Gi

resources:
  requests:
    cpu: 250m
    memory: 512Mi
  limits:
    cpu: "1"
    memory: 1Gi

redis:
  configmap: |
    maxmemory 768mb
    maxmemory-policy allkeys-lru
    cluster-require-full-coverage no
    cluster-allow-reads-when-down yes

metrics:
  enabled: true
  serviceMonitor:
    enabled: true

podAntiAffinityPreset: hard
# Deploy Redis Cluster:
helm install redis-cluster bitnami/redis-cluster \
  --namespace caching \
  -f redis-cluster-values.yaml

# Verify cluster:
kubectl -n caching exec redis-cluster-0 -- \
  redis-cli -a "$REDIS_PASSWORD" cluster info
# cluster_state:ok
# cluster_slots_assigned:16384
# cluster_slots_ok:16384
# cluster_size:3
# cluster_known_nodes:6

# Check slot distribution:
kubectl -n caching exec redis-cluster-0 -- \
  redis-cli -a "$REDIS_PASSWORD" cluster nodes
# node-id master 10.244.1.x:6379 0-5460
# node-id master 10.244.2.x:6379 5461-10922
# node-id master 10.244.3.x:6379 10923-16383

PART 4: PERSISTENCE — RDB vs AOF

FeatureRDB (Snapshotting)AOF (Append-Only File)RDB + AOF
MechanismPoint-in-time snapshotLog every write operationBoth
Data Loss_Up to last snapshot~1 second (appendfsync everysec)Minimal
Recovery SpeedFast (load binary)Slower (replay operations)Uses RDB first
Disk I/OPeriodic burst (fork)Continuous (small writes)Both
File SizeCompact_Larger (rewrite helps)Both files
Best ForCaching (acceptable loss)Session store (minimal loss)Production (recommended)
# Configuration cho production (RDB + AOF):
# save 900 1       → Snapshot nếu 1 key change trong 900s
# save 300 10      → Snapshot nếu 10 keys change trong 300s  
# save 60 10000    → Snapshot nếu 10000 keys change trong 60s
# appendonly yes
# appendfsync everysec

PART 5: CACHING STRATEGIES


1. Cache-Aside (Lazy Loading):
   App ──► Cache HIT? ──► Return data
                │ MISS
                ▼
           Read from DB → Store in Cache → Return

2. Write-Through:
   App ──► Write to Cache ──► Write to DB (sync)

3. Write-Behind (Write-Back):
   App ──► Write to Cache ──► Async write to DB

4. Read-Through:
   App ──► Cache (auto-loads from DB on miss)
# Python example - Cache-Aside pattern:
import redis
import json

r = redis.Redis(
    host='redis-sentinel.caching.svc',
    port=26379,
    password=os.environ['REDIS_PASSWORD'],
    sentinel_manager=True,
    db=0,
    decode_responses=True
)

def get_user(user_id):
    cache_key = f"user:{user_id}"
    
    # Try cache first:
    cached = r.get(cache_key)
    if cached:
        return json.loads(cached)
    
    # Cache miss → query DB:
    user = db.query("SELECT * FROM users WHERE id = %s", user_id)
    
    # Store in cache with TTL:
    r.setex(cache_key, 3600, json.dumps(user))  # 1 hour TTL
    
    return user

def update_user(user_id, data):
    # Update DB first:
    db.execute("UPDATE users SET ... WHERE id = %s", user_id)
    
    # Invalidate cache:
    r.delete(f"user:{user_id}")

PART 6: APPLICATION CONNECTION

# App connecting to Redis Sentinel:
apiVersion: apps/v1
kind: Deployment
metadata:
  name: my-app
spec:
  template:
    spec:
      containers:
        - name: app
          env:
            - name: REDIS_SENTINEL_HOST
              value: "redis-sentinel.caching.svc"
            - name: REDIS_SENTINEL_PORT
              value: "26379"
            - name: REDIS_MASTER_NAME
              value: "mymaster"
            - name: REDIS_PASSWORD
              valueFrom:
                secretKeyRef:
                  name: redis-app-secret
                  key: password

PART 7: MONITORING

# Key Redis metrics:
# redis_connected_clients          — Current connections
# redis_used_memory_bytes          — Memory usage
# redis_evicted_keys_total         — Keys evicted (maxmemory)
# redis_keyspace_hits_total        — Cache hits
# redis_keyspace_misses_total      — Cache misses
# redis_commands_processed_total   — Commands/sec

# Cache Hit Rate formula:
# hit_rate = hits / (hits + misses) × 100%
# Target: > 90%

# Grafana dashboard: ID 11835 (Redis Dashboard for Prometheus)
# Alerting:
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
  name: redis-alerts
  namespace: caching
spec:
  groups:
    - name: redis
      rules:
        - alert: RedisHighMemory
          expr: redis_used_memory_bytes / redis_maxmemory > 0.9
          for: 5m
          labels:
            severity: warning

        - alert: RedisMasterDown
          expr: redis_up{role="master"} == 0
          for: 1m
          labels:
            severity: critical

        - alert: RedisLowCacheHitRate
          expr: |
            redis_keyspace_hits_total / 
            (redis_keyspace_hits_total + redis_keyspace_misses_total) < 0.8
          for: 15m
          labels:
            severity: warning

💡 KEY TAKEAWAYS

  1. Sentinel: Simple HA, single master, good for < 32GB datasets
  2. Cluster: Horizontal scaling, multi-master, for large datasets
  3. Persistence: RDB + AOF for production, balance durability vs performance
  4. maxmemory-policy: most popular allkeys-lru for caching
  5. Cache-Aside: Recommended pattern, invalidate on write
  6. Monitor: Cache hit rate > 90%, memory usage, evictions

🎯 EXERCISE

Exercise 1: Redis Sentinel Failover

  • Deploy Redis Sentinel 3-node
  • Write data to master, kill master pod
  • Verify Sentinel promotes replica, data intact__HTMLTAG_306___

Exercise 2: Benchmark

  • Run redis-benchmark: SET/GET 100,000 keys
  • Compare latency: with/without persistence__HTMLTAG_314___
  • Monitor memory fragmentation ratio

📚 NEXT POST

In Lesson 24: Istio Service Mesh Architecture, we will learn service mesh and deploy Istio for microservices communication.