Chuyển đến nội dung chính

LESSON 19: DAEMONSETS

DaemonSets ensure one Pod runs on each node. Use cases: monitoring agent, logging agent, network plugin, storage. Real-life example deploying Grafana Alloy collector on every node.

🔒 DevSecOps — Lesson 19 LESSON 19: DAEMONSETS

KUBERNETES: FROM BASIC TO ADVANCED

Module 5: Workload Management__HTMLTAG_60___

xdev.asia

DaemonSets — One Pod Per Node__HTMLTAG_66___

When you need a Pod running on all nodes in the cluster (or a specific subset of nodes), DaemonSet is the answer. Unlike Deployment, which distributes Pods in replicas, DaemonSet ensures that each node (matching the selector) always has its own Pod. When a new node is added to the cluster, DaemonSet automatically creates a Pod on that node. When the node is deleted, the Pod is also garbage collected.

1. What is DaemonSet? One Pod Per Node Mechanism

DaemonSet controller continuously ensures:

  • Each matching node has exactly one DaemonSet Pod running
  • When a new node joins the cluster → A new Pod is automatically created
  • When node is removed from cluster → Pod is deleted
  • Deleting DaemonSet will clean up all Pods it created

The most basic DaemonSet:

apiVersion: apps/v1
kind: DaemonSet
metadata:
  name: simple-daemon
  namespace: monitoring
  labels:
    app: simple-daemon
spec:
  selector:
    matchLabels:
      name: simple-daemon
  template:
    metadata:
      labels:
        name: simple-daemon
    spec:
      containers:
      - name: daemon-container
        image: busybox:1.35
        command: ["sh", "-c", "while true; do echo $(hostname) is alive; sleep 60; done"]
        resources:
          requests:
            cpu: "10m"
            memory: "16Mi"
          limits:
            cpu: "50m"
            memory: "32Mi"

2. Actual Use Cases

DaemonSet is widely used for infrastructure agents:

2.1 Logging Agents__HTMLTAG_94___
  • Grafana Alloy: Collect logs from stdout files and containers on each node
  • Fluent Bit: Lightweight log forwarder, forward logs to Elasticsearch/Loki
  • Fluentd: Aggregate and transform logs before shipping

2.2 Monitoring Agents

  • Prometheus Node Exporter: Export CPU, memory, disk, network metrics of node
  • Datadog Agent: Full observability — metrics, logs, traces, processes
  • Elastic Agent: Unified Elastic Stack agent

2.3 Network Plugins (CNI)

  • Cilium: eBPF-based networking, security, observability
  • Calico: Network policy enforcement
  • Weave Net: Simple overlay network

2.4 Storage (CSI Node Plugin)

  • AWS EBS CSI Driver node plugin: Mount EBS volumes on EC2 nodes
  • Longhorn: Distributed block storage — engine runs on each node
  • OpenEBS: Cloud-native storage

2.5 Security

  • Falco: Runtime threat detection using eBPF (see lesson 26)
  • NeuVector: Container security platform
  • Sysdig Agent: Security and performance monitoring

3. DaemonSet vs Deployment: When to Use Which?

FAQ: "Why not use Deployment with replicas: N instead of DaemonSet?"

Use DaemonSet when:

  • Need direct access to node resources (filesystem, network interfaces, hardware)
  • Need Pods on each specific node__HTMLTAG_188___, not any N pods__HTMLTAG_189___
  • Agent needs to know exactly which node it is running on (hostname, node IP)
  • Infrastructure agents: logging, monitoring, CNI, CSI

Use Deployment when:

  • Need to scale the number of replicas independently of the number of nodes
  • Stateless application without node-specific access
  • Web servers, API servers, common microservices

4. Node Selection — Select Nodes for DaemonSet

DaemonSet does not always need to run across the entire cluster. You can limit it in many ways.

4.1 nodeSelector__HTMLTAG_212___

Only deploy on nodes with a specific label:

spec:
  template:
    spec:
      nodeSelector:
        node-role: worker          # Chỉ chạy trên worker nodes
        disk-type: ssd             # Chỉ trên nodes có SSD

4.2 Node Affinity

More complex than nodeSelector, supports operators like In, NotIn, Exists:

spec:
  template:
    spec:
      affinity:
        nodeAffinity:
          requiredDuringSchedulingIgnoredDuringExecution:
            nodeSelectorTerms:
            - matchExpressions:
              - key: kubernetes.io/os
                operator: In
                values:
                - linux
              - key: node.kubernetes.io/instance-type
                operator: NotIn
                values:
                - t3.nano
                - t3.micro      # Bỏ qua nodes quá nhỏ

4.3 Tolerations — Deploy to Tainted Nodes

Nodes can have taints to prevent regular Pods from scheduling there. DaemonSet often needs tolerations to bypass these taints — especially to run on control plane nodes:

spec:
  template:
    spec:
      tolerations:
      # Tolerate taint trên master/control-plane nodes
      - key: node-role.kubernetes.io/control-plane
        operator: Exists
        effect: NoSchedule
      # Tolerate node đang không ready
      - key: node.kubernetes.io/not-ready
        operator: Exists
        effect: NoExecute
        tolerationSeconds: 300
      # Tolerate node unreachable
      - key: node.kubernetes.io/unreachable
        operator: Exists
        effect: NoExecute
        tolerationSeconds: 300

The last two tolerations (not-ready and unreachable) are very important for infrastructure agents — you want logging/monitoring to continue running even when the node is having problems, to capture logs of that problem.

5. Update Strategies

DaemonSet has two update strategies:

5.1 RollingUpdate (Default)

spec:
  updateStrategy:
    type: RollingUpdate
    rollingUpdate:
      maxUnavailable: 1   # Tối đa 1 Pod không available cùng lúc
      # Hoặc dùng percentage:
      # maxUnavailable: 10%

With RollingUpdate, Kubernetes updates each Pod one by one (or according to maxUnavailable). Old Pod is deleted, new Pod is created on the same node.

5.2 OnDelete

spec:
  updateStrategy:
    type: OnDelete

With OnDelete, DaemonSet will not automatically update Pods. New Pod (with new spec) is only created when you manually delete the old Pod. Use when you want complete control over the update process.

6. Practical Example: Grafana Alloy Node Agent

Grafana Alloy is an OpenTelemetry-native collector from Grafana Labs, replacing the old Grafana Agent. Deploy Alloy as DaemonSet to collect logs and metrics from each node:

apiVersion: apps/v1
kind: DaemonSet
metadata:
  name: grafana-alloy
  namespace: monitoring
  labels:
    app.kubernetes.io/name: alloy
    app.kubernetes.io/component: logs-collector
spec:
  selector:
    matchLabels:
      app.kubernetes.io/name: alloy
  updateStrategy:
    type: RollingUpdate
    rollingUpdate:
      maxUnavailable: 1
  template:
    metadata:
      labels:
        app.kubernetes.io/name: alloy
    spec:
      serviceAccountName: grafana-alloy
      tolerations:
      - key: node-role.kubernetes.io/control-plane
        operator: Exists
        effect: NoSchedule
      - key: node.kubernetes.io/not-ready
        operator: Exists
        effect: NoExecute
        tolerationSeconds: 300
      - key: node.kubernetes.io/unreachable
        operator: Exists
        effect: NoExecute
        tolerationSeconds: 300
      # Alloy cần chạy với host network để collect system metrics
      hostNetwork: false
      hostPID: false
      containers:
      - name: alloy
        image: grafana/alloy:v1.1.1
        args:
        - run
        - /etc/alloy/config.alloy
        - --storage.path=/tmp/alloy
        - --server.http.listen-addr=0.0.0.0:12345
        env:
        - name: HOSTNAME
          valueFrom:
            fieldRef:
              fieldPath: spec.nodeName
        - name: LOKI_URL
          value: "http://loki-gateway.monitoring.svc.cluster.local/loki/api/v1/push"
        ports:
        - name: http-metrics
          containerPort: 12345
          protocol: TCP
        resources:
          requests:
            cpu: "50m"
            memory: "128Mi"
          limits:
            cpu: "200m"
            memory: "256Mi"
        volumeMounts:
        - name: config
          mountPath: /etc/alloy
        - name: varlog
          mountPath: /var/log
          readOnly: true
        - name: varlibdockercontainers
          mountPath: /var/lib/docker/containers
          readOnly: true
        - name: storage
          mountPath: /tmp/alloy
        securityContext:
          allowPrivilegeEscalation: false
          readOnlyRootFilesystem: true
          runAsNonRoot: false   # Cần root để đọc /var/log
          runAsUser: 0
          capabilities:
            drop:
            - ALL
            add:
            - DAC_READ_SEARCH  # Đọc files của user khác
      volumes:
      - name: config
        configMap:
          name: grafana-alloy-config
      - name: varlog
        hostPath:
          path: /var/log
      - name: varlibdockercontainers
        hostPath:
          path: /var/lib/docker/containers
      - name: storage
        emptyDir: {}

ConfigMap for Alloy config:

apiVersion: v1
kind: ConfigMap
metadata:
  name: grafana-alloy-config
  namespace: monitoring
data:
  config.alloy: |
    // Collect container logs
    discovery.kubernetes "pods" {
      role = "pod"
      namespaces {
        own_namespace = false
      }
    }

    discovery.relabel "pod_logs" {
      targets = discovery.kubernetes.pods.targets

      rule {
        source_labels = ["__meta_kubernetes_namespace"]
        action = "replace"
        target_label = "namespace"
      }

      rule {
        source_labels = ["__meta_kubernetes_pod_name"]
        action = "replace"
        target_label = "pod"
      }

      rule {
        source_labels = ["__meta_kubernetes_container_name"]
        action = "replace"
        target_label = "container"
      }

      rule {
        source_labels = ["__meta_kubernetes_node_name"]
        target_label = "node"
      }
    }

    loki.source.kubernetes "pod_logs" {
      targets    = discovery.relabel.pod_logs.output
      forward_to = [loki.write.default.receiver]
    }

    loki.write "default" {
      endpoint {
        url = env("LOKI_URL")
      }
    }

7. Example: Node Exporter for Prometheus Metrics

Node Exporter collects node hardware and OS metrics, exposing them to Prometheus scrape:

apiVersion: apps/v1
kind: DaemonSet
metadata:
  name: node-exporter
  namespace: monitoring
  labels:
    app.kubernetes.io/name: node-exporter
    app.kubernetes.io/component: metrics
spec:
  selector:
    matchLabels:
      app.kubernetes.io/name: node-exporter
  updateStrategy:
    type: RollingUpdate
    rollingUpdate:
      maxUnavailable: 1
  template:
    metadata:
      labels:
        app.kubernetes.io/name: node-exporter
      annotations:
        prometheus.io/scrape: "true"
        prometheus.io/port: "9100"
        prometheus.io/path: "/metrics"
    spec:
      hostPID: true       # Cần để monitor processes
      hostIPC: true
      hostNetwork: true   # Cần để collect network metrics chính xác
      dnsPolicy: ClusterFirstWithHostNet
      tolerations:
      - operator: Exists    # Tolerate tất cả taints
      serviceAccountName: node-exporter
      containers:
      - name: node-exporter
        image: quay.io/prometheus/node-exporter:v1.7.0
        args:
        - --path.rootfs=/host/root
        - --path.procfs=/host/proc
        - --path.sysfs=/host/sys
        - --collector.filesystem.mount-points-exclude=^/(dev|proc|sys|var/lib/docker/.+|var/lib/kubelet/.+)($|/)
        - --no-collector.ipvs
        ports:
        - containerPort: 9100
          hostPort: 9100    # Expose trực tiếp trên node
          name: metrics
          protocol: TCP
        resources:
          requests:
            cpu: "15m"
            memory: "32Mi"
          limits:
            cpu: "250m"
            memory: "180Mi"
        securityContext:
          readOnlyRootFilesystem: true
          runAsNonRoot: true
          runAsUser: 65534    # nobody
          allowPrivilegeEscalation: false
        volumeMounts:
        - name: root
          mountPath: /host/root
          readOnly: true
          mountPropagation: HostToContainer
        - name: proc
          mountPath: /host/proc
          readOnly: true
        - name: sys
          mountPath: /host/sys
          readOnly: true
      volumes:
      - name: root
        hostPath:
          path: /
      - name: proc
        hostPath:
          path: /proc
      - name: sys
        hostPath:
          path: /sys

8. Priority When Deploy DaemonSet: PriorityClass

Infrastructure agents like logging and monitoring need to be scheduled before application Pods — even when the node is under memory pressure. Use PriorityClass to ensure this:

apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
  name: system-node-agent
value: 1000000      # Cao hơn workloads thông thường (default: 0)
globalDefault: false
description: "Priority class for node-level DaemonSet agents"
preemptionPolicy: PreemptLowerPriority

Then assign PriorityClass to DaemonSet:

spec:
  template:
    spec:
      priorityClassName: system-node-agent
      containers:
      - name: node-agent
        # ...

Kubernetes has a number of built-in priority classes:

  • system-cluster-critical: 2000000000 — for CoreDNS, kube-proxy
  • system-node-critical: 2000001000 — for kubelet-adjacent components

9. Test and Debug DaemonSet

# Xem tất cả DaemonSets trong cluster
kubectl get daemonsets --all-namespaces

# Xem chi tiết DaemonSet
kubectl describe daemonset node-exporter -n monitoring

# Kiểm tra số Pods expected vs actual
# DESIRED: số node phù hợp | CURRENT: đang create | READY: healthy | UP-TO-DATE: có spec mới | AVAILABLE: ready for use
kubectl get daemonset grafana-alloy -n monitoring
# NAME            DESIRED   CURRENT   READY   UP-TO-DATE   AVAILABLE
# grafana-alloy   5         5         5       5            5

# Xem Pods của DaemonSet trên từng node
kubectl get pods -n monitoring -l app.kubernetes.io/name=alloy -o wide

# Debug Pod không schedule
kubectl get events -n monitoring --sort-by='.lastTimestamp'

# Xem logs từ một node cụ thể
kubectl logs -n monitoring -l app.kubernetes.io/name=alloy \
  --field-selector spec.nodeName=worker-node-1

10. DaemonSet Best Practices

  • Reasonable Resources: DaemonSet runs on every node — if each Pod uses 500m of CPU, the whole cluster loses a significant amount of capacity__HTMLTAG_285___
  • Full Tolerations: Add tolerations for not-ready and unreachable so that the agent continues to operate when the node has problems__HTMLTAG_293___
  • PriorityClass: Assign priority higher than regular workloads to ensure schedule
  • readOnlyRootFilesystem: Enable when possible, use emptyDir or hostPath for data to write__HTMLTAG_301___
  • RollingUpdate maxUnavailable: Keep to 1 or a small number to not lose coverage on many nodes at the same time__HTMLTAG_305___
  • Namespace isolation: Deploy DaemonSet into separate namespace (e.g., monitoring, logging) — do not mix with application workloads

DaemonSet is an indispensable tool in every Kubernetes cluster production. Understanding how to properly use DaemonSet helps you build a solid infrastructure observability and security layer on the Kubernetes platform.