Chuyển đến nội dung chính

BÀI 19: DAEMONSETS

DaemonSets đảm bảo một Pod chạy trên mỗi node. Use cases: monitoring agent, logging agent, network plugin, storage. Ví dụ thực tế deploy Grafana Alloy collector trên mọi node.

🔒 DevSecOps — Bài 19 BÀI 19: DAEMONSETS

KUBERNETES: TỪ CƠ BẢN ĐẾN NÂNG CAO

Module 5: Workload Management

xdev.asia

DaemonSets — Một Pod Trên Mỗi Node

Khi bạn cần một Pod chạy trên tất cả các nodes trong cluster (hoặc một subset nodes cụ thể), DaemonSet là câu trả lời. Không giống Deployment phân phối Pods theo replicas, DaemonSet đảm bảo mỗi node (phù hợp với selector) luôn có đúng một Pod của nó. Khi node mới được thêm vào cluster, DaemonSet tự động tạo Pod trên node đó. Khi node bị xóa, Pod cũng được garbage collected.

1. DaemonSet là Gì? Cơ Chế Một Pod Per Node

DaemonSet controller liên tục đảm bảo:

  • Mỗi node phù hợp có đúng một Pod của DaemonSet đang chạy
  • Khi node mới join cluster → Pod mới được tạo tự động
  • Khi node bị remove khỏi cluster → Pod bị xóa
  • Xóa DaemonSet sẽ dọn sạch tất cả Pods nó tạo ra

DaemonSet cơ bản nhất:

apiVersion: apps/v1
kind: DaemonSet
metadata:
  name: simple-daemon
  namespace: monitoring
  labels:
    app: simple-daemon
spec:
  selector:
    matchLabels:
      name: simple-daemon
  template:
    metadata:
      labels:
        name: simple-daemon
    spec:
      containers:
      - name: daemon-container
        image: busybox:1.35
        command: ["sh", "-c", "while true; do echo $(hostname) is alive; sleep 60; done"]
        resources:
          requests:
            cpu: "10m"
            memory: "16Mi"
          limits:
            cpu: "50m"
            memory: "32Mi"

2. Use Cases Thực Tế

DaemonSet được dùng rộng rãi cho các loại infrastructure agents:

2.1 Logging Agents

  • Grafana Alloy: Thu thập logs từ files và container stdout trên mỗi node
  • Fluent Bit: Lightweight log forwarder, forward logs đến Elasticsearch/Loki
  • Fluentd: Aggregate và transform logs trước khi ship đi

2.2 Monitoring Agents

  • Prometheus Node Exporter: Export CPU, memory, disk, network metrics của node
  • Datadog Agent: Full observability — metrics, logs, traces, processes
  • Elastic Agent: Unified Elastic Stack agent

2.3 Network Plugins (CNI)

  • Cilium: eBPF-based networking, security, observability
  • Calico: Network policy enforcement
  • Weave Net: Simple overlay network

2.4 Storage (CSI Node Plugin)

  • AWS EBS CSI Driver node plugin: Mount EBS volumes trên EC2 nodes
  • Longhorn: Distributed block storage — engine chạy trên mỗi node
  • OpenEBS: Cloud-native storage

2.5 Security

  • Falco: Runtime threat detection bằng eBPF (xem bài 26)
  • NeuVector: Container security platform
  • Sysdig Agent: Security và performance monitoring

3. DaemonSet vs Deployment: Khi Nào Dùng Gì?

Câu hỏi thường gặp: "Tại sao không dùng Deployment với replicas: N thay vì DaemonSet?"

Dùng DaemonSet khi:

  • Cần access trực tiếp đến node resources (filesystem, network interfaces, hardware)
  • Cần Pod trên mỗi node cụ thể, không phải N pods bất kỳ
  • Agent cần biết chính xác nó đang chạy trên node nào (hostname, node IP)
  • Infrastructure agents: logging, monitoring, CNI, CSI

Dùng Deployment khi:

  • Cần scale số lượng replica độc lập với số node
  • Application stateless không cần node-specific access
  • Web servers, API servers, microservices thông thường

4. Node Selection — Chọn Nodes Cho DaemonSet

DaemonSet không phải lúc nào cũng cần chạy trên toàn bộ cluster. Bạn có thể giới hạn bằng nhiều cách.

4.1 nodeSelector

Chỉ deploy trên nodes có label cụ thể:

spec:
  template:
    spec:
      nodeSelector:
        node-role: worker          # Chỉ chạy trên worker nodes
        disk-type: ssd             # Chỉ trên nodes có SSD

4.2 Node Affinity

Phức tạp hơn nodeSelector, hỗ trợ operators như In, NotIn, Exists:

spec:
  template:
    spec:
      affinity:
        nodeAffinity:
          requiredDuringSchedulingIgnoredDuringExecution:
            nodeSelectorTerms:
            - matchExpressions:
              - key: kubernetes.io/os
                operator: In
                values:
                - linux
              - key: node.kubernetes.io/instance-type
                operator: NotIn
                values:
                - t3.nano
                - t3.micro      # Bỏ qua nodes quá nhỏ

4.3 Tolerations — Deploy Lên Tainted Nodes

Nodes có thể có taints để ngăn Pods thông thường schedule lên đó. DaemonSet thường cần tolerations để bypass các taints này — đặc biệt là để chạy trên control plane nodes:

spec:
  template:
    spec:
      tolerations:
      # Tolerate taint trên master/control-plane nodes
      - key: node-role.kubernetes.io/control-plane
        operator: Exists
        effect: NoSchedule
      # Tolerate node đang không ready
      - key: node.kubernetes.io/not-ready
        operator: Exists
        effect: NoExecute
        tolerationSeconds: 300
      # Tolerate node unreachable
      - key: node.kubernetes.io/unreachable
        operator: Exists
        effect: NoExecute
        tolerationSeconds: 300

Hai toleration cuối (not-ready và unreachable) rất quan trọng cho infrastructure agents — bạn muốn logging/monitoring tiếp tục chạy kể cả khi node đang có vấn đề, để capture logs về sự cố đó.

5. Update Strategies

DaemonSet có hai update strategies:

5.1 RollingUpdate (Default)

spec:
  updateStrategy:
    type: RollingUpdate
    rollingUpdate:
      maxUnavailable: 1   # Tối đa 1 Pod không available cùng lúc
      # Hoặc dùng percentage:
      # maxUnavailable: 10%

Với RollingUpdate, Kubernetes update từng Pod một (hoặc theo maxUnavailable). Pod cũ bị xóa, Pod mới được tạo trên cùng node.

5.2 OnDelete

spec:
  updateStrategy:
    type: OnDelete

Với OnDelete, DaemonSet sẽ không tự update Pods. Pod mới (với spec mới) chỉ được tạo khi bạn manually xóa Pod cũ. Dùng khi bạn muốn kiểm soát hoàn toàn quá trình update.

6. Ví Dụ Thực Tế: Grafana Alloy Node Agent

Grafana Alloy là OpenTelemetry-native collector của Grafana Labs, thay thế Grafana Agent cũ. Deploy Alloy như DaemonSet để thu thập logs và metrics từ mỗi node:

apiVersion: apps/v1
kind: DaemonSet
metadata:
  name: grafana-alloy
  namespace: monitoring
  labels:
    app.kubernetes.io/name: alloy
    app.kubernetes.io/component: logs-collector
spec:
  selector:
    matchLabels:
      app.kubernetes.io/name: alloy
  updateStrategy:
    type: RollingUpdate
    rollingUpdate:
      maxUnavailable: 1
  template:
    metadata:
      labels:
        app.kubernetes.io/name: alloy
    spec:
      serviceAccountName: grafana-alloy
      tolerations:
      - key: node-role.kubernetes.io/control-plane
        operator: Exists
        effect: NoSchedule
      - key: node.kubernetes.io/not-ready
        operator: Exists
        effect: NoExecute
        tolerationSeconds: 300
      - key: node.kubernetes.io/unreachable
        operator: Exists
        effect: NoExecute
        tolerationSeconds: 300
      # Alloy cần chạy với host network để collect system metrics
      hostNetwork: false
      hostPID: false
      containers:
      - name: alloy
        image: grafana/alloy:v1.1.1
        args:
        - run
        - /etc/alloy/config.alloy
        - --storage.path=/tmp/alloy
        - --server.http.listen-addr=0.0.0.0:12345
        env:
        - name: HOSTNAME
          valueFrom:
            fieldRef:
              fieldPath: spec.nodeName
        - name: LOKI_URL
          value: "http://loki-gateway.monitoring.svc.cluster.local/loki/api/v1/push"
        ports:
        - name: http-metrics
          containerPort: 12345
          protocol: TCP
        resources:
          requests:
            cpu: "50m"
            memory: "128Mi"
          limits:
            cpu: "200m"
            memory: "256Mi"
        volumeMounts:
        - name: config
          mountPath: /etc/alloy
        - name: varlog
          mountPath: /var/log
          readOnly: true
        - name: varlibdockercontainers
          mountPath: /var/lib/docker/containers
          readOnly: true
        - name: storage
          mountPath: /tmp/alloy
        securityContext:
          allowPrivilegeEscalation: false
          readOnlyRootFilesystem: true
          runAsNonRoot: false   # Cần root để đọc /var/log
          runAsUser: 0
          capabilities:
            drop:
            - ALL
            add:
            - DAC_READ_SEARCH  # Đọc files của user khác
      volumes:
      - name: config
        configMap:
          name: grafana-alloy-config
      - name: varlog
        hostPath:
          path: /var/log
      - name: varlibdockercontainers
        hostPath:
          path: /var/lib/docker/containers
      - name: storage
        emptyDir: {}

ConfigMap cho Alloy config:

apiVersion: v1
kind: ConfigMap
metadata:
  name: grafana-alloy-config
  namespace: monitoring
data:
  config.alloy: |
    // Collect container logs
    discovery.kubernetes "pods" {
      role = "pod"
      namespaces {
        own_namespace = false
      }
    }

    discovery.relabel "pod_logs" {
      targets = discovery.kubernetes.pods.targets

      rule {
        source_labels = ["__meta_kubernetes_namespace"]
        action = "replace"
        target_label = "namespace"
      }

      rule {
        source_labels = ["__meta_kubernetes_pod_name"]
        action = "replace"
        target_label = "pod"
      }

      rule {
        source_labels = ["__meta_kubernetes_container_name"]
        action = "replace"
        target_label = "container"
      }

      rule {
        source_labels = ["__meta_kubernetes_node_name"]
        target_label = "node"
      }
    }

    loki.source.kubernetes "pod_logs" {
      targets    = discovery.relabel.pod_logs.output
      forward_to = [loki.write.default.receiver]
    }

    loki.write "default" {
      endpoint {
        url = env("LOKI_URL")
      }
    }

7. Ví Dụ: Node Exporter cho Prometheus Metrics

Node Exporter thu thập metrics phần cứng và OS của node, expose cho Prometheus scrape:

apiVersion: apps/v1
kind: DaemonSet
metadata:
  name: node-exporter
  namespace: monitoring
  labels:
    app.kubernetes.io/name: node-exporter
    app.kubernetes.io/component: metrics
spec:
  selector:
    matchLabels:
      app.kubernetes.io/name: node-exporter
  updateStrategy:
    type: RollingUpdate
    rollingUpdate:
      maxUnavailable: 1
  template:
    metadata:
      labels:
        app.kubernetes.io/name: node-exporter
      annotations:
        prometheus.io/scrape: "true"
        prometheus.io/port: "9100"
        prometheus.io/path: "/metrics"
    spec:
      hostPID: true       # Cần để monitor processes
      hostIPC: true
      hostNetwork: true   # Cần để collect network metrics chính xác
      dnsPolicy: ClusterFirstWithHostNet
      tolerations:
      - operator: Exists    # Tolerate tất cả taints
      serviceAccountName: node-exporter
      containers:
      - name: node-exporter
        image: quay.io/prometheus/node-exporter:v1.7.0
        args:
        - --path.rootfs=/host/root
        - --path.procfs=/host/proc
        - --path.sysfs=/host/sys
        - --collector.filesystem.mount-points-exclude=^/(dev|proc|sys|var/lib/docker/.+|var/lib/kubelet/.+)($|/)
        - --no-collector.ipvs
        ports:
        - containerPort: 9100
          hostPort: 9100    # Expose trực tiếp trên node
          name: metrics
          protocol: TCP
        resources:
          requests:
            cpu: "15m"
            memory: "32Mi"
          limits:
            cpu: "250m"
            memory: "180Mi"
        securityContext:
          readOnlyRootFilesystem: true
          runAsNonRoot: true
          runAsUser: 65534    # nobody
          allowPrivilegeEscalation: false
        volumeMounts:
        - name: root
          mountPath: /host/root
          readOnly: true
          mountPropagation: HostToContainer
        - name: proc
          mountPath: /host/proc
          readOnly: true
        - name: sys
          mountPath: /host/sys
          readOnly: true
      volumes:
      - name: root
        hostPath:
          path: /
      - name: proc
        hostPath:
          path: /proc
      - name: sys
        hostPath:
          path: /sys

8. Ưu Tiên Khi Deploy DaemonSet: PriorityClass

Infrastructure agents như logging và monitoring cần được schedule trước khi application Pods — ngay cả khi node đang chịu áp lực memory. Dùng PriorityClass để đảm bảo điều này:

apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
  name: system-node-agent
value: 1000000      # Cao hơn workloads thông thường (default: 0)
globalDefault: false
description: "Priority class for node-level DaemonSet agents"
preemptionPolicy: PreemptLowerPriority

Sau đó assign PriorityClass cho DaemonSet:

spec:
  template:
    spec:
      priorityClassName: system-node-agent
      containers:
      - name: node-agent
        # ...

Kubernetes có một số built-in priority classes:

  • system-cluster-critical: 2000000000 — cho CoreDNS, kube-proxy
  • system-node-critical: 2000001000 — cho kubelet-adjacent components

9. Kiểm Tra và Debug DaemonSet

# Xem tất cả DaemonSets trong cluster
kubectl get daemonsets --all-namespaces

# Xem chi tiết DaemonSet
kubectl describe daemonset node-exporter -n monitoring

# Kiểm tra số Pods expected vs actual
# DESIRED: số node phù hợp | CURRENT: đang create | READY: healthy | UP-TO-DATE: có spec mới | AVAILABLE: ready for use
kubectl get daemonset grafana-alloy -n monitoring
# NAME            DESIRED   CURRENT   READY   UP-TO-DATE   AVAILABLE
# grafana-alloy   5         5         5       5            5

# Xem Pods của DaemonSet trên từng node
kubectl get pods -n monitoring -l app.kubernetes.io/name=alloy -o wide

# Debug Pod không schedule
kubectl get events -n monitoring --sort-by='.lastTimestamp'

# Xem logs từ một node cụ thể
kubectl logs -n monitoring -l app.kubernetes.io/name=alloy \
  --field-selector spec.nodeName=worker-node-1

10. DaemonSet Best Practices

  • Resources hợp lý: DaemonSet chạy trên mọi node — nếu mỗi Pod dùng 500m CPU, cả cluster mất đi đáng kể capacity
  • Tolerations đầy đủ: Thêm tolerations cho not-ready và unreachable để agent tiếp tục hoạt động khi node có vấn đề
  • PriorityClass: Assign priority cao hơn workloads thông thường để đảm bảo schedule
  • readOnlyRootFilesystem: Bật khi có thể, dùng emptyDir hoặc hostPath cho data cần write
  • RollingUpdate maxUnavailable: Giữ ở 1 hoặc một con số nhỏ để không mất coverage trên nhiều nodes cùng lúc
  • Namespace isolation: Deploy DaemonSet vào namespace riêng (e.g., monitoring, logging) — không mix với application workloads

DaemonSet là công cụ không thể thiếu trong mọi Kubernetes cluster production. Hiểu cách dùng đúng DaemonSet giúp bạn xây dựng infrastructure observability và security layer vững chắc trên nền tảng Kubernetes.