Chuyển đến nội dung chính

Bài 7: SecurityContext, Capabilities & ServiceAccounts

SecurityContext cho Pod và Container: runAsUser, runAsNonRoot, readOnlyRootFilesystem, Linux capabilities. ServiceAccount binding và automountServiceAccountToken.

SecurityContext — Pod-level vs Container-level, Linux capabilities

1. SecurityContext

SecurityContext định nghĩa các privilege và access control settings cho Pod hoặc Container.

apiVersion: v1
kind: Pod
spec:
  securityContext:            # Pod-level: applies to ALL containers
    runAsUser: 1000           # UID để chạy containers
    runAsGroup: 3000          # GID primary group
    fsGroup: 2000             # GID cho mounted volumes
    runAsNonRoot: true        # Reject containers that run as root

  containers:
  - name: app
    image: myapp
    securityContext:          # Container-level: overrides pod-level
      allowPrivilegeEscalation: false
      readOnlyRootFilesystem: true   # Filesystem là read-only
      capabilities:
        add: ["NET_BIND_SERVICE"]    # Add capability
        drop: ["ALL"]               # Drop all, add only what needed
SettingLevelTác dụng
runAsUserPod/ContainerChạy process với UID cụ thể
runAsNonRootPod/ContainerTừ chối chạy nếu UID = 0 (root)
readOnlyRootFilesystemContainerMount root filesystem read-only
allowPrivilegeEscalationContainerChặn privilege escalation (sudo etc.)
privilegedContainerRun as privileged (như root trên host)
fsGroupPodGID cho volume files (shared volume access)

Exam tip: Container-level securityContext override Pod-level settings. Nếu Pod có runAsUser: 1000 và container có runAsUser: 2000, container đó chạy với UID 2000. Hay bị test: verify user bằng kubectl exec pod -- id hoặc whoami.

2. Linux Capabilities

Capabilities cho phép grant specific privileges mà không cần full root.

# Ví dụ thường gặp:
NET_BIND_SERVICE  — Bind to port < 1024 (e.g., port 80)
NET_ADMIN         — Network administration (ifconfig etc.)
SYS_TIME          — Modify system clock
CHOWN             — Change file ownership
SETUID/SETGID     — Change user/group ID

securityContext:
  capabilities:
    drop: ["ALL"]             # Best practice: drop all first
    add: ["NET_BIND_SERVICE"] # Then re-add only what's needed

3. ServiceAccounts

Pod dùng ServiceAccount để authenticate với Kubernetes API.

# Tạo ServiceAccount
kubectl create serviceaccount my-sa

# Bind vào Role
kubectl create rolebinding my-binding \
  --role=pod-reader \
  --serviceaccount=default:my-sa

# Assign SA cho Pod
spec:
  serviceAccountName: my-sa     # Use specific SA
  automountServiceAccountToken: false  # Don't auto-mount SA token

# Mặc định: default SA được mount tại /var/run/secrets/kubernetes.io/serviceaccount/
# Token file có thể gọi K8s API từ container
ConceptMô tả
Default SAMỗi namespace có sẵn default SA (ít quyền)
Token mountToken tự động mount vào pod nếu không tắt
automountServiceAccountToken: falseTắt việc mount token (best security practice)

Exam tip: Khi Pod cần gọi Kubernetes API (ví dụ: operator pattern), cần ServiceAccount có quyền phù hợp. Nếu Pod không cần gọi API, best practice là set automountServiceAccountToken: false. CKAD thường test: tạo SA, bind role, và set SA trong Pod spec.

4. readOnlyRootFilesystem với emptyDir

# Khi dùng readOnlyRootFilesystem: true, app KHÔNG ghi vào root FS.
# Nhưng app vẫn cần ghi temp files → dùng emptyDir volume:

spec:
  containers:
  - name: app
    image: myapp
    securityContext:
      readOnlyRootFilesystem: true
    volumeMounts:
    - name: tmp
      mountPath: /tmp        # App ghi temp files vào đây
    - name: cache
      mountPath: /app/cache
  volumes:
  - name: tmp
    emptyDir: {}
  - name: cache
    emptyDir: {}

5. Cheat Sheet

TaskYAML / Command
Run container as non-rootsecurityContext: runAsNonRoot: true
Run với specific UIDsecurityContext: runAsUser: 1000
Read-only filesystemsecurityContext: readOnlyRootFilesystem: true
Drop all capabilitiescapabilities: drop: ["ALL"]
Gán ServiceAccountspec: serviceAccountName: my-sa
Verify user trong containerkubectl exec pod -- whoami

6. Practice Questions

Q1: A Pod spec has securityContext.runAsUser: 1000 at the Pod level. One container within the Pod has securityContext.runAsUser: 2000. What UID does that container run with?

  • A) 0 (root, because Pod-level overrides)
  • B) 1000 (Pod-level takes priority)
  • C) 2000 (Container-level overrides Pod-level) ✓
  • D) Both UIDs simultaneously

Explanation: Container-level securityContext settings override Pod-level settings. The container runs with UID 2000. Other containers in the same Pod without a container-level securityContext.runAsUser would inherit the Pod-level UID of 1000.

Q2: An application container needs to bind to port 80 (a privileged port below 1024) but should NOT run as root. How do you configure this?

  • A) Set securityContext.privileged: true
  • B) Set securityContext.runAsUser: 0
  • C) Add NET_BIND_SERVICE capability while dropping ALL others ✓
  • D) Use a NodePort Service instead of port 80

Explanation: Linux capabilities allow granular privilege grants. NET_BIND_SERVICE allows binding to ports below 1024 without full root. Best practice is to drop ALL capabilities first, then add only what's needed: capabilities: { drop: ["ALL"], add: ["NET_BIND_SERVICE"] }.

Q3: A Pod is running with readOnlyRootFilesystem: true, but the application tries to write to /tmp and fails. What is the best solution?

  • A) Remove readOnlyRootFilesystem: true
  • B) Set securityContext.privileged: true
  • C) Mount an emptyDir volume at /tmp ✓
  • D) Use a ConfigMap mounted at /tmp

Explanation: readOnlyRootFilesystem prevents writes to the container's filesystem, but emptyDir volumes are separate writable mounts. By mounting an emptyDir at /tmp, the application can write temp files there while the root filesystem remains read-only — maintaining the security benefit.