DaemonSets — One Pod Per Node__HTMLTAG_66___
When you need a Pod running on all nodes in the cluster (or a specific subset of nodes), DaemonSet is the answer. Unlike Deployment, which distributes Pods in replicas, DaemonSet ensures that each node (matching the selector) always has its own Pod. When a new node is added to the cluster, DaemonSet automatically creates a Pod on that node. When the node is deleted, the Pod is also garbage collected.
1. What is DaemonSet? One Pod Per Node Mechanism
DaemonSet controller continuously ensures:
- Each matching node has exactly one DaemonSet Pod running
- When a new node joins the cluster → A new Pod is automatically created
- When node is removed from cluster → Pod is deleted
- Deleting DaemonSet will clean up all Pods it created
The most basic DaemonSet:
apiVersion: apps/v1
kind: DaemonSet
metadata:
name: simple-daemon
namespace: monitoring
labels:
app: simple-daemon
spec:
selector:
matchLabels:
name: simple-daemon
template:
metadata:
labels:
name: simple-daemon
spec:
containers:
- name: daemon-container
image: busybox:1.35
command: ["sh", "-c", "while true; do echo $(hostname) is alive; sleep 60; done"]
resources:
requests:
cpu: "10m"
memory: "16Mi"
limits:
cpu: "50m"
memory: "32Mi"
2. Actual Use Cases
DaemonSet is widely used for infrastructure agents:
2.1 Logging Agents__HTMLTAG_94___
- Grafana Alloy: Collect logs from stdout files and containers on each node
- Fluent Bit: Lightweight log forwarder, forward logs to Elasticsearch/Loki
- Fluentd: Aggregate and transform logs before shipping
2.2 Monitoring Agents
- Prometheus Node Exporter: Export CPU, memory, disk, network metrics of node
- Datadog Agent: Full observability — metrics, logs, traces, processes
- Elastic Agent: Unified Elastic Stack agent
2.3 Network Plugins (CNI)
- Cilium: eBPF-based networking, security, observability
- Calico: Network policy enforcement
- Weave Net: Simple overlay network
2.4 Storage (CSI Node Plugin)
- AWS EBS CSI Driver node plugin: Mount EBS volumes on EC2 nodes
- Longhorn: Distributed block storage — engine runs on each node
- OpenEBS: Cloud-native storage
2.5 Security
- Falco: Runtime threat detection using eBPF (see lesson 26)
- NeuVector: Container security platform
- Sysdig Agent: Security and performance monitoring
3. DaemonSet vs Deployment: When to Use Which?
FAQ: "Why not use Deployment with replicas: N instead of DaemonSet?"
Use DaemonSet when:
- Need direct access to node resources (filesystem, network interfaces, hardware)
- Need Pods on each specific node__HTMLTAG_188___, not any N pods__HTMLTAG_189___
- Agent needs to know exactly which node it is running on (hostname, node IP)
- Infrastructure agents: logging, monitoring, CNI, CSI
Use Deployment when:
- Need to scale the number of replicas independently of the number of nodes
- Stateless application without node-specific access
- Web servers, API servers, common microservices
4. Node Selection — Select Nodes for DaemonSet
DaemonSet does not always need to run across the entire cluster. You can limit it in many ways.
4.1 nodeSelector__HTMLTAG_212___
Only deploy on nodes with a specific label:
spec:
template:
spec:
nodeSelector:
node-role: worker # Chỉ chạy trên worker nodes
disk-type: ssd # Chỉ trên nodes có SSD
4.2 Node Affinity
More complex than nodeSelector, supports operators like In, NotIn, Exists:
spec:
template:
spec:
affinity:
nodeAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
nodeSelectorTerms:
- matchExpressions:
- key: kubernetes.io/os
operator: In
values:
- linux
- key: node.kubernetes.io/instance-type
operator: NotIn
values:
- t3.nano
- t3.micro # Bỏ qua nodes quá nhỏ
4.3 Tolerations — Deploy to Tainted Nodes
Nodes can have taints to prevent regular Pods from scheduling there. DaemonSet often needs tolerations to bypass these taints — especially to run on control plane nodes:
spec:
template:
spec:
tolerations:
# Tolerate taint trên master/control-plane nodes
- key: node-role.kubernetes.io/control-plane
operator: Exists
effect: NoSchedule
# Tolerate node đang không ready
- key: node.kubernetes.io/not-ready
operator: Exists
effect: NoExecute
tolerationSeconds: 300
# Tolerate node unreachable
- key: node.kubernetes.io/unreachable
operator: Exists
effect: NoExecute
tolerationSeconds: 300
The last two tolerations (not-ready and unreachable) are very important for infrastructure agents — you want logging/monitoring to continue running even when the node is having problems, to capture logs of that problem.
5. Update Strategies
DaemonSet has two update strategies:
5.1 RollingUpdate (Default)
spec:
updateStrategy:
type: RollingUpdate
rollingUpdate:
maxUnavailable: 1 # Tối đa 1 Pod không available cùng lúc
# Hoặc dùng percentage:
# maxUnavailable: 10%
With RollingUpdate, Kubernetes updates each Pod one by one (or according to maxUnavailable). Old Pod is deleted, new Pod is created on the same node.
5.2 OnDelete
spec:
updateStrategy:
type: OnDelete
With OnDelete, DaemonSet will not automatically update Pods. New Pod (with new spec) is only created when you manually delete the old Pod. Use when you want complete control over the update process.
6. Practical Example: Grafana Alloy Node Agent
Grafana Alloy is an OpenTelemetry-native collector from Grafana Labs, replacing the old Grafana Agent. Deploy Alloy as DaemonSet to collect logs and metrics from each node:
apiVersion: apps/v1
kind: DaemonSet
metadata:
name: grafana-alloy
namespace: monitoring
labels:
app.kubernetes.io/name: alloy
app.kubernetes.io/component: logs-collector
spec:
selector:
matchLabels:
app.kubernetes.io/name: alloy
updateStrategy:
type: RollingUpdate
rollingUpdate:
maxUnavailable: 1
template:
metadata:
labels:
app.kubernetes.io/name: alloy
spec:
serviceAccountName: grafana-alloy
tolerations:
- key: node-role.kubernetes.io/control-plane
operator: Exists
effect: NoSchedule
- key: node.kubernetes.io/not-ready
operator: Exists
effect: NoExecute
tolerationSeconds: 300
- key: node.kubernetes.io/unreachable
operator: Exists
effect: NoExecute
tolerationSeconds: 300
# Alloy cần chạy với host network để collect system metrics
hostNetwork: false
hostPID: false
containers:
- name: alloy
image: grafana/alloy:v1.1.1
args:
- run
- /etc/alloy/config.alloy
- --storage.path=/tmp/alloy
- --server.http.listen-addr=0.0.0.0:12345
env:
- name: HOSTNAME
valueFrom:
fieldRef:
fieldPath: spec.nodeName
- name: LOKI_URL
value: "http://loki-gateway.monitoring.svc.cluster.local/loki/api/v1/push"
ports:
- name: http-metrics
containerPort: 12345
protocol: TCP
resources:
requests:
cpu: "50m"
memory: "128Mi"
limits:
cpu: "200m"
memory: "256Mi"
volumeMounts:
- name: config
mountPath: /etc/alloy
- name: varlog
mountPath: /var/log
readOnly: true
- name: varlibdockercontainers
mountPath: /var/lib/docker/containers
readOnly: true
- name: storage
mountPath: /tmp/alloy
securityContext:
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
runAsNonRoot: false # Cần root để đọc /var/log
runAsUser: 0
capabilities:
drop:
- ALL
add:
- DAC_READ_SEARCH # Đọc files của user khác
volumes:
- name: config
configMap:
name: grafana-alloy-config
- name: varlog
hostPath:
path: /var/log
- name: varlibdockercontainers
hostPath:
path: /var/lib/docker/containers
- name: storage
emptyDir: {}
ConfigMap for Alloy config:
apiVersion: v1
kind: ConfigMap
metadata:
name: grafana-alloy-config
namespace: monitoring
data:
config.alloy: |
// Collect container logs
discovery.kubernetes "pods" {
role = "pod"
namespaces {
own_namespace = false
}
}
discovery.relabel "pod_logs" {
targets = discovery.kubernetes.pods.targets
rule {
source_labels = ["__meta_kubernetes_namespace"]
action = "replace"
target_label = "namespace"
}
rule {
source_labels = ["__meta_kubernetes_pod_name"]
action = "replace"
target_label = "pod"
}
rule {
source_labels = ["__meta_kubernetes_container_name"]
action = "replace"
target_label = "container"
}
rule {
source_labels = ["__meta_kubernetes_node_name"]
target_label = "node"
}
}
loki.source.kubernetes "pod_logs" {
targets = discovery.relabel.pod_logs.output
forward_to = [loki.write.default.receiver]
}
loki.write "default" {
endpoint {
url = env("LOKI_URL")
}
}
7. Example: Node Exporter for Prometheus Metrics
Node Exporter collects node hardware and OS metrics, exposing them to Prometheus scrape:
apiVersion: apps/v1
kind: DaemonSet
metadata:
name: node-exporter
namespace: monitoring
labels:
app.kubernetes.io/name: node-exporter
app.kubernetes.io/component: metrics
spec:
selector:
matchLabels:
app.kubernetes.io/name: node-exporter
updateStrategy:
type: RollingUpdate
rollingUpdate:
maxUnavailable: 1
template:
metadata:
labels:
app.kubernetes.io/name: node-exporter
annotations:
prometheus.io/scrape: "true"
prometheus.io/port: "9100"
prometheus.io/path: "/metrics"
spec:
hostPID: true # Cần để monitor processes
hostIPC: true
hostNetwork: true # Cần để collect network metrics chính xác
dnsPolicy: ClusterFirstWithHostNet
tolerations:
- operator: Exists # Tolerate tất cả taints
serviceAccountName: node-exporter
containers:
- name: node-exporter
image: quay.io/prometheus/node-exporter:v1.7.0
args:
- --path.rootfs=/host/root
- --path.procfs=/host/proc
- --path.sysfs=/host/sys
- --collector.filesystem.mount-points-exclude=^/(dev|proc|sys|var/lib/docker/.+|var/lib/kubelet/.+)($|/)
- --no-collector.ipvs
ports:
- containerPort: 9100
hostPort: 9100 # Expose trực tiếp trên node
name: metrics
protocol: TCP
resources:
requests:
cpu: "15m"
memory: "32Mi"
limits:
cpu: "250m"
memory: "180Mi"
securityContext:
readOnlyRootFilesystem: true
runAsNonRoot: true
runAsUser: 65534 # nobody
allowPrivilegeEscalation: false
volumeMounts:
- name: root
mountPath: /host/root
readOnly: true
mountPropagation: HostToContainer
- name: proc
mountPath: /host/proc
readOnly: true
- name: sys
mountPath: /host/sys
readOnly: true
volumes:
- name: root
hostPath:
path: /
- name: proc
hostPath:
path: /proc
- name: sys
hostPath:
path: /sys
8. Priority When Deploy DaemonSet: PriorityClass
Infrastructure agents like logging and monitoring need to be scheduled before application Pods — even when the node is under memory pressure. Use PriorityClass to ensure this:
apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
name: system-node-agent
value: 1000000 # Cao hơn workloads thông thường (default: 0)
globalDefault: false
description: "Priority class for node-level DaemonSet agents"
preemptionPolicy: PreemptLowerPriority
Then assign PriorityClass to DaemonSet:
spec:
template:
spec:
priorityClassName: system-node-agent
containers:
- name: node-agent
# ...
Kubernetes has a number of built-in priority classes:
system-cluster-critical: 2000000000 — for CoreDNS, kube-proxysystem-node-critical: 2000001000 — for kubelet-adjacent components
9. Test and Debug DaemonSet
# Xem tất cả DaemonSets trong cluster
kubectl get daemonsets --all-namespaces
# Xem chi tiết DaemonSet
kubectl describe daemonset node-exporter -n monitoring
# Kiểm tra số Pods expected vs actual
# DESIRED: số node phù hợp | CURRENT: đang create | READY: healthy | UP-TO-DATE: có spec mới | AVAILABLE: ready for use
kubectl get daemonset grafana-alloy -n monitoring
# NAME DESIRED CURRENT READY UP-TO-DATE AVAILABLE
# grafana-alloy 5 5 5 5 5
# Xem Pods của DaemonSet trên từng node
kubectl get pods -n monitoring -l app.kubernetes.io/name=alloy -o wide
# Debug Pod không schedule
kubectl get events -n monitoring --sort-by='.lastTimestamp'
# Xem logs từ một node cụ thể
kubectl logs -n monitoring -l app.kubernetes.io/name=alloy \
--field-selector spec.nodeName=worker-node-1
10. DaemonSet Best Practices
- Reasonable Resources: DaemonSet runs on every node — if each Pod uses 500m of CPU, the whole cluster loses a significant amount of capacity__HTMLTAG_285___
- Full Tolerations: Add tolerations for
not-readyandunreachableso that the agent continues to operate when the node has problems__HTMLTAG_293___ - PriorityClass: Assign priority higher than regular workloads to ensure schedule
- readOnlyRootFilesystem: Enable when possible, use emptyDir or hostPath for data to write__HTMLTAG_301___
- RollingUpdate maxUnavailable: Keep to 1 or a small number to not lose coverage on many nodes at the same time__HTMLTAG_305___
- Namespace isolation: Deploy DaemonSet into separate namespace (e.g.,
monitoring,logging) — do not mix with application workloads
DaemonSet is an indispensable tool in every Kubernetes cluster production. Understanding how to properly use DaemonSet helps you build a solid infrastructure observability and security layer on the Kubernetes platform.