DaemonSets — Một Pod Trên Mỗi Node
Khi bạn cần một Pod chạy trên tất cả các nodes trong cluster (hoặc một subset nodes cụ thể), DaemonSet là câu trả lời. Không giống Deployment phân phối Pods theo replicas, DaemonSet đảm bảo mỗi node (phù hợp với selector) luôn có đúng một Pod của nó. Khi node mới được thêm vào cluster, DaemonSet tự động tạo Pod trên node đó. Khi node bị xóa, Pod cũng được garbage collected.
1. DaemonSet là Gì? Cơ Chế Một Pod Per Node
DaemonSet controller liên tục đảm bảo:
- Mỗi node phù hợp có đúng một Pod của DaemonSet đang chạy
- Khi node mới join cluster → Pod mới được tạo tự động
- Khi node bị remove khỏi cluster → Pod bị xóa
- Xóa DaemonSet sẽ dọn sạch tất cả Pods nó tạo ra
DaemonSet cơ bản nhất:
apiVersion: apps/v1
kind: DaemonSet
metadata:
name: simple-daemon
namespace: monitoring
labels:
app: simple-daemon
spec:
selector:
matchLabels:
name: simple-daemon
template:
metadata:
labels:
name: simple-daemon
spec:
containers:
- name: daemon-container
image: busybox:1.35
command: ["sh", "-c", "while true; do echo $(hostname) is alive; sleep 60; done"]
resources:
requests:
cpu: "10m"
memory: "16Mi"
limits:
cpu: "50m"
memory: "32Mi"
2. Use Cases Thực Tế
DaemonSet được dùng rộng rãi cho các loại infrastructure agents:
2.1 Logging Agents
- Grafana Alloy: Thu thập logs từ files và container stdout trên mỗi node
- Fluent Bit: Lightweight log forwarder, forward logs đến Elasticsearch/Loki
- Fluentd: Aggregate và transform logs trước khi ship đi
2.2 Monitoring Agents
- Prometheus Node Exporter: Export CPU, memory, disk, network metrics của node
- Datadog Agent: Full observability — metrics, logs, traces, processes
- Elastic Agent: Unified Elastic Stack agent
2.3 Network Plugins (CNI)
- Cilium: eBPF-based networking, security, observability
- Calico: Network policy enforcement
- Weave Net: Simple overlay network
2.4 Storage (CSI Node Plugin)
- AWS EBS CSI Driver node plugin: Mount EBS volumes trên EC2 nodes
- Longhorn: Distributed block storage — engine chạy trên mỗi node
- OpenEBS: Cloud-native storage
2.5 Security
- Falco: Runtime threat detection bằng eBPF (xem bài 26)
- NeuVector: Container security platform
- Sysdig Agent: Security và performance monitoring
3. DaemonSet vs Deployment: Khi Nào Dùng Gì?
Câu hỏi thường gặp: "Tại sao không dùng Deployment với replicas: N thay vì DaemonSet?"
Dùng DaemonSet khi:
- Cần access trực tiếp đến node resources (filesystem, network interfaces, hardware)
- Cần Pod trên mỗi node cụ thể, không phải N pods bất kỳ
- Agent cần biết chính xác nó đang chạy trên node nào (hostname, node IP)
- Infrastructure agents: logging, monitoring, CNI, CSI
Dùng Deployment khi:
- Cần scale số lượng replica độc lập với số node
- Application stateless không cần node-specific access
- Web servers, API servers, microservices thông thường
4. Node Selection — Chọn Nodes Cho DaemonSet
DaemonSet không phải lúc nào cũng cần chạy trên toàn bộ cluster. Bạn có thể giới hạn bằng nhiều cách.
4.1 nodeSelector
Chỉ deploy trên nodes có label cụ thể:
spec:
template:
spec:
nodeSelector:
node-role: worker # Chỉ chạy trên worker nodes
disk-type: ssd # Chỉ trên nodes có SSD
4.2 Node Affinity
Phức tạp hơn nodeSelector, hỗ trợ operators như In, NotIn, Exists:
spec:
template:
spec:
affinity:
nodeAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
nodeSelectorTerms:
- matchExpressions:
- key: kubernetes.io/os
operator: In
values:
- linux
- key: node.kubernetes.io/instance-type
operator: NotIn
values:
- t3.nano
- t3.micro # Bỏ qua nodes quá nhỏ
4.3 Tolerations — Deploy Lên Tainted Nodes
Nodes có thể có taints để ngăn Pods thông thường schedule lên đó. DaemonSet thường cần tolerations để bypass các taints này — đặc biệt là để chạy trên control plane nodes:
spec:
template:
spec:
tolerations:
# Tolerate taint trên master/control-plane nodes
- key: node-role.kubernetes.io/control-plane
operator: Exists
effect: NoSchedule
# Tolerate node đang không ready
- key: node.kubernetes.io/not-ready
operator: Exists
effect: NoExecute
tolerationSeconds: 300
# Tolerate node unreachable
- key: node.kubernetes.io/unreachable
operator: Exists
effect: NoExecute
tolerationSeconds: 300
Hai toleration cuối (not-ready và unreachable) rất quan trọng cho infrastructure agents — bạn muốn logging/monitoring tiếp tục chạy kể cả khi node đang có vấn đề, để capture logs về sự cố đó.
5. Update Strategies
DaemonSet có hai update strategies:
5.1 RollingUpdate (Default)
spec:
updateStrategy:
type: RollingUpdate
rollingUpdate:
maxUnavailable: 1 # Tối đa 1 Pod không available cùng lúc
# Hoặc dùng percentage:
# maxUnavailable: 10%
Với RollingUpdate, Kubernetes update từng Pod một (hoặc theo maxUnavailable). Pod cũ bị xóa, Pod mới được tạo trên cùng node.
5.2 OnDelete
spec:
updateStrategy:
type: OnDelete
Với OnDelete, DaemonSet sẽ không tự update Pods. Pod mới (với spec mới) chỉ được tạo khi bạn manually xóa Pod cũ. Dùng khi bạn muốn kiểm soát hoàn toàn quá trình update.
6. Ví Dụ Thực Tế: Grafana Alloy Node Agent
Grafana Alloy là OpenTelemetry-native collector của Grafana Labs, thay thế Grafana Agent cũ. Deploy Alloy như DaemonSet để thu thập logs và metrics từ mỗi node:
apiVersion: apps/v1
kind: DaemonSet
metadata:
name: grafana-alloy
namespace: monitoring
labels:
app.kubernetes.io/name: alloy
app.kubernetes.io/component: logs-collector
spec:
selector:
matchLabels:
app.kubernetes.io/name: alloy
updateStrategy:
type: RollingUpdate
rollingUpdate:
maxUnavailable: 1
template:
metadata:
labels:
app.kubernetes.io/name: alloy
spec:
serviceAccountName: grafana-alloy
tolerations:
- key: node-role.kubernetes.io/control-plane
operator: Exists
effect: NoSchedule
- key: node.kubernetes.io/not-ready
operator: Exists
effect: NoExecute
tolerationSeconds: 300
- key: node.kubernetes.io/unreachable
operator: Exists
effect: NoExecute
tolerationSeconds: 300
# Alloy cần chạy với host network để collect system metrics
hostNetwork: false
hostPID: false
containers:
- name: alloy
image: grafana/alloy:v1.1.1
args:
- run
- /etc/alloy/config.alloy
- --storage.path=/tmp/alloy
- --server.http.listen-addr=0.0.0.0:12345
env:
- name: HOSTNAME
valueFrom:
fieldRef:
fieldPath: spec.nodeName
- name: LOKI_URL
value: "http://loki-gateway.monitoring.svc.cluster.local/loki/api/v1/push"
ports:
- name: http-metrics
containerPort: 12345
protocol: TCP
resources:
requests:
cpu: "50m"
memory: "128Mi"
limits:
cpu: "200m"
memory: "256Mi"
volumeMounts:
- name: config
mountPath: /etc/alloy
- name: varlog
mountPath: /var/log
readOnly: true
- name: varlibdockercontainers
mountPath: /var/lib/docker/containers
readOnly: true
- name: storage
mountPath: /tmp/alloy
securityContext:
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
runAsNonRoot: false # Cần root để đọc /var/log
runAsUser: 0
capabilities:
drop:
- ALL
add:
- DAC_READ_SEARCH # Đọc files của user khác
volumes:
- name: config
configMap:
name: grafana-alloy-config
- name: varlog
hostPath:
path: /var/log
- name: varlibdockercontainers
hostPath:
path: /var/lib/docker/containers
- name: storage
emptyDir: {}
ConfigMap cho Alloy config:
apiVersion: v1
kind: ConfigMap
metadata:
name: grafana-alloy-config
namespace: monitoring
data:
config.alloy: |
// Collect container logs
discovery.kubernetes "pods" {
role = "pod"
namespaces {
own_namespace = false
}
}
discovery.relabel "pod_logs" {
targets = discovery.kubernetes.pods.targets
rule {
source_labels = ["__meta_kubernetes_namespace"]
action = "replace"
target_label = "namespace"
}
rule {
source_labels = ["__meta_kubernetes_pod_name"]
action = "replace"
target_label = "pod"
}
rule {
source_labels = ["__meta_kubernetes_container_name"]
action = "replace"
target_label = "container"
}
rule {
source_labels = ["__meta_kubernetes_node_name"]
target_label = "node"
}
}
loki.source.kubernetes "pod_logs" {
targets = discovery.relabel.pod_logs.output
forward_to = [loki.write.default.receiver]
}
loki.write "default" {
endpoint {
url = env("LOKI_URL")
}
}
7. Ví Dụ: Node Exporter cho Prometheus Metrics
Node Exporter thu thập metrics phần cứng và OS của node, expose cho Prometheus scrape:
apiVersion: apps/v1
kind: DaemonSet
metadata:
name: node-exporter
namespace: monitoring
labels:
app.kubernetes.io/name: node-exporter
app.kubernetes.io/component: metrics
spec:
selector:
matchLabels:
app.kubernetes.io/name: node-exporter
updateStrategy:
type: RollingUpdate
rollingUpdate:
maxUnavailable: 1
template:
metadata:
labels:
app.kubernetes.io/name: node-exporter
annotations:
prometheus.io/scrape: "true"
prometheus.io/port: "9100"
prometheus.io/path: "/metrics"
spec:
hostPID: true # Cần để monitor processes
hostIPC: true
hostNetwork: true # Cần để collect network metrics chính xác
dnsPolicy: ClusterFirstWithHostNet
tolerations:
- operator: Exists # Tolerate tất cả taints
serviceAccountName: node-exporter
containers:
- name: node-exporter
image: quay.io/prometheus/node-exporter:v1.7.0
args:
- --path.rootfs=/host/root
- --path.procfs=/host/proc
- --path.sysfs=/host/sys
- --collector.filesystem.mount-points-exclude=^/(dev|proc|sys|var/lib/docker/.+|var/lib/kubelet/.+)($|/)
- --no-collector.ipvs
ports:
- containerPort: 9100
hostPort: 9100 # Expose trực tiếp trên node
name: metrics
protocol: TCP
resources:
requests:
cpu: "15m"
memory: "32Mi"
limits:
cpu: "250m"
memory: "180Mi"
securityContext:
readOnlyRootFilesystem: true
runAsNonRoot: true
runAsUser: 65534 # nobody
allowPrivilegeEscalation: false
volumeMounts:
- name: root
mountPath: /host/root
readOnly: true
mountPropagation: HostToContainer
- name: proc
mountPath: /host/proc
readOnly: true
- name: sys
mountPath: /host/sys
readOnly: true
volumes:
- name: root
hostPath:
path: /
- name: proc
hostPath:
path: /proc
- name: sys
hostPath:
path: /sys
8. Ưu Tiên Khi Deploy DaemonSet: PriorityClass
Infrastructure agents như logging và monitoring cần được schedule trước khi application Pods — ngay cả khi node đang chịu áp lực memory. Dùng PriorityClass để đảm bảo điều này:
apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
name: system-node-agent
value: 1000000 # Cao hơn workloads thông thường (default: 0)
globalDefault: false
description: "Priority class for node-level DaemonSet agents"
preemptionPolicy: PreemptLowerPriority
Sau đó assign PriorityClass cho DaemonSet:
spec:
template:
spec:
priorityClassName: system-node-agent
containers:
- name: node-agent
# ...
Kubernetes có một số built-in priority classes:
system-cluster-critical: 2000000000 — cho CoreDNS, kube-proxysystem-node-critical: 2000001000 — cho kubelet-adjacent components
9. Kiểm Tra và Debug DaemonSet
# Xem tất cả DaemonSets trong cluster
kubectl get daemonsets --all-namespaces
# Xem chi tiết DaemonSet
kubectl describe daemonset node-exporter -n monitoring
# Kiểm tra số Pods expected vs actual
# DESIRED: số node phù hợp | CURRENT: đang create | READY: healthy | UP-TO-DATE: có spec mới | AVAILABLE: ready for use
kubectl get daemonset grafana-alloy -n monitoring
# NAME DESIRED CURRENT READY UP-TO-DATE AVAILABLE
# grafana-alloy 5 5 5 5 5
# Xem Pods của DaemonSet trên từng node
kubectl get pods -n monitoring -l app.kubernetes.io/name=alloy -o wide
# Debug Pod không schedule
kubectl get events -n monitoring --sort-by='.lastTimestamp'
# Xem logs từ một node cụ thể
kubectl logs -n monitoring -l app.kubernetes.io/name=alloy \
--field-selector spec.nodeName=worker-node-1
10. DaemonSet Best Practices
- Resources hợp lý: DaemonSet chạy trên mọi node — nếu mỗi Pod dùng 500m CPU, cả cluster mất đi đáng kể capacity
- Tolerations đầy đủ: Thêm tolerations cho
not-readyvàunreachableđể agent tiếp tục hoạt động khi node có vấn đề - PriorityClass: Assign priority cao hơn workloads thông thường để đảm bảo schedule
- readOnlyRootFilesystem: Bật khi có thể, dùng emptyDir hoặc hostPath cho data cần write
- RollingUpdate maxUnavailable: Giữ ở 1 hoặc một con số nhỏ để không mất coverage trên nhiều nodes cùng lúc
- Namespace isolation: Deploy DaemonSet vào namespace riêng (e.g.,
monitoring,logging) — không mix với application workloads
DaemonSet là công cụ không thể thiếu trong mọi Kubernetes cluster production. Hiểu cách dùng đúng DaemonSet giúp bạn xây dựng infrastructure observability và security layer vững chắc trên nền tảng Kubernetes.