Chuyển đến nội dung chính

LESSON 11: PERSISTENT STORAGE AND CSI

Manage storage with PersistentVolumes, PersistentVolumeClaims, StorageClasses. CSI drivers must replace in-tree plugins (removed K8s 1.31). Dynamic provisioning, volume snapshots, new VolumeAttributesClass (K8s 1.29+).

🔒 DevSecOps — Lesson 11 LESSON 11: PERSISTENT STORAGE AND CSI

KUBERNETES: FROM BASIC TO ADVANCED

Module 3: Configuration & Storage

xdev.asia

Persistent Storage and CSI in Kubernetes__HTMLTAG_66___

Containers are ephemeral — when a container restarts or Pod is rescheduled on another node, all data inside is lost. This is a serious problem with stateful workloads like databases, message queues, and file storage. Kubernetes solves this problem with a rich storage system, and from Kubernetes 1.30-1.31, in-tree storage plugins have been completely removed, replaced by Container Storage Interface (CSI) standardized drivers.

Storage Lifecycle: From Ephemeral to Persistent__HTMLTAG_72___

emptyDir: Ephemeral Storage

emptyDir creates an empty directory when the Pod is created and persists for the life of the Pod. When the Pod is deleted, the data is completely lost. Suitable for caches, temporary files, sharing data between containers in the same Pod.

apiVersion: v1
kind: Pod
metadata:
  name: shared-storage-pod
spec:
  containers:
  - name: writer
    image: busybox
    command: ["/bin/sh", "-c"]
    args: ["while true; do date >> /shared/log.txt; sleep 1; done"]
    volumeMounts:
    - name: shared-data
      mountPath: /shared
  - name: reader
    image: busybox
    command: ["/bin/sh", "-c"]
    args: ["tail -f /shared/log.txt"]
    volumeMounts:
    - name: shared-data
      mountPath: /shared
  volumes:
  - name: shared-data
    emptyDir:
      sizeLimit: 500Mi  # Giới hạn kích thước
      medium: Memory    # Lưu trong RAM (tmpfs) - nhanh hơn nhưng dùng memory

hostPath: Node-Local Storage

hostPath mounts a path on the host node to the container. The data exists when the container restarts but is lost when the Pod is scheduled to another node. Should not be used in production unless there is a special reason (system daemons, monitoring agents).

apiVersion: v1
kind: Pod
metadata:
  name: host-path-pod
spec:
  containers:
  - name: app
    image: nginx:1.27
    volumeMounts:
    - name: host-volume
      mountPath: /data
  volumes:
  - name: host-volume
    hostPath:
      path: /data/myapp
      type: DirectoryOrCreate  # Tạo directory nếu chưa tồn tại

PersistentVolume and PersistentVolumeClaim

Core Concepts

Kubernetes separates providing storage (admin) and using storage (developer) through two abstractions:

  • PersistentVolume (PV): Cluster-level storage resource created by admin or dynamically provisioned. Represents a piece of actual storage (NFS share, cloud disk, etc.)
  • PersistentVolumeClaim (PVC): User request for storage. Developers just need to declare "I need 10Gi of storage with ReadWriteOnce" without knowing where the storage is.

PersistentVolume Definition

apiVersion: v1
kind: PersistentVolume
metadata:
  name: pv-database-001
  labels:
    type: ssd
    environment: production
spec:
  capacity:
    storage: 100Gi
  volumeMode: Filesystem  # hoặc Block
  accessModes:
  - ReadWriteOnce  # Chỉ một node mount được tại một thời điểm
  persistentVolumeReclaimPolicy: Retain  # Giữ lại data sau khi PVC released
  storageClassName: premium-ssd
  # CSI volume source (thay vì in-tree)
  csi:
    driver: ebs.csi.aws.com
    volumeHandle: vol-0a1b2c3d4e5f67890
    fsType: ext4
    volumeAttributes:
      storage.kubernetes.io/csiProvisionerIdentity: "1234567890"

PersistentVolumeClaim

apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: database-storage
  namespace: production
spec:
  accessModes:
  - ReadWriteOnce
  volumeMode: Filesystem
  resources:
    requests:
      storage: 50Gi
  storageClassName: premium-ssd
  # Selector để bind với specific PV (optional)
  selector:
    matchLabels:
      type: ssd
      environment: production
# Xem trạng thái PVC
kubectl get pvc -n production

# Output:
# NAME               STATUS   VOLUME              CAPACITY   ACCESS MODES   STORAGECLASS   AGE
# database-storage   Bound    pv-database-001     100Gi      RWO            premium-ssd    5m

Using PVC in Pod

apiVersion: apps/v1
kind: Deployment
metadata:
  name: postgres
  namespace: production
spec:
  replicas: 1
  selector:
    matchLabels:
      app: postgres
  template:
    metadata:
      labels:
        app: postgres
    spec:
      containers:
      - name: postgres
        image: postgres:16
        env:
        - name: POSTGRES_PASSWORD
          valueFrom:
            secretKeyRef:
              name: postgres-secret
              key: password
        - name: PGDATA
          value: /var/lib/postgresql/data/pgdata
        volumeMounts:
        - name: postgres-storage
          mountPath: /var/lib/postgresql/data
      volumes:
      - name: postgres-storage
        persistentVolumeClaim:
          claimName: database-storage

Access Modes

Kubernetes defines 4 access modes, showing how volumes can be mounted:

  • ReadWriteOnce (RWO): A node that can mount read-write. Most common with block storage (EBS, GCE PD). From K8s 1.22+, RWO allows multiple Pods on the SAME read-write node.
  • ReadOnlyMany (ROX): Multiple nodes can mount read-only at the same time. Matches shared config/data.
  • ReadWriteMany (RWX): Many nodes mount read-write. Requires network filesystem such as NFS, CephFS, Azure Files. Important: block storage (EBS, GCE PD) does NOT support RWX.
  • ReadWriteOncePod (RWOP): Only a single Pod in the entire cluster can mount. Stronger than RWO — ensures exclusive access at Pod level, not just node level. Requires CSI driver support.

StorageClass: Dynamic Provisioning

Why StorageClass?

Instead of admins having to create PVs manually, StorageClass allows dynamic provisioning — Kubernetes automatically creates PVs when there is a PVC request. Admin only needs to define storage "class" (eg: ssd-fast, hdd-cheap, nfs-shared).

# StorageClass với AWS EBS CSI Driver
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
  name: ebs-gp3
  annotations:
    storageclass.kubernetes.io/is-default-class: "true"
provisioner: ebs.csi.aws.com
parameters:
  type: gp3
  iops: "3000"
  throughput: "125"
  encrypted: "true"
  kmsKeyId: arn:aws:kms:ap-southeast-1:123456789012:key/mrk-abc123
volumeBindingMode: WaitForFirstConsumer  # Đợi Pod được schedule trước khi tạo volume
reclaimPolicy: Delete  # Xóa PV khi PVC bị xóa
allowVolumeExpansion: true  # Cho phép resize PVC
---
# StorageClass cho Longhorn
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
  name: longhorn-replicated
provisioner: driver.longhorn.io
parameters:
  numberOfReplicas: "3"
  staleReplicaTimeout: "2880"
  fromBackup: ""
  fsType: "ext4"
volumeBindingMode: Immediate
reclaimPolicy: Delete
allowVolumeExpansion: true

Reclaim Policy

  • Delete: PV and underlying storage are deleted when PVC is deleted. Default with dynamic provisioning. Suitable for ephemeral workloads.
  • Retain: The PV is retained (Released state) when the PVC is deleted. Admin must reclaim manually. Suitable for production databases that need data protection.
  • Recycle: Deprecated, not recommended.

CSI: Container Storage Interface

Why Was In-Tree Plugins Removed?

Previously, Kubernetes had many in-tree storage plugins compiled directly into the core Kubernetes binary (aws-ebs, gce-pd, azure-disk, cephfs, nfs...). This creates many problems:

  • Bug in storage plugin can crash entire kube-apiserver/kubelet
  • Release cycle of storage drivers tied to Kubernetes release
  • Difficult to maintain when the number of plugins increases
  • Storage vendors cannot ship fixes independently

CSI (Container Storage Interface) solves all of these problems by standardizing the interface between Kubernetes and storage providers. Drivers run as separate Pods, can be updated independently.

Timeline Remove In-Tree Plugins__HTMLTAG_180___
  • K8s 1.26-1.28: Many in-tree plugins deprecated
  • K8s 1.29: In-tree NFS and many plugins converted to deprecated, CSI required
  • K8s 1.30: In-tree NFS plugin removed
  • K8s 1.31: In-tree CephFS, Ceph RBD plugins completely removed__HTMLTAG_197___
  • K8s 1.32+: Continue removing the remaining in-tree plugins

Popular CSI Drivers

Cloud Providers

# AWS EBS CSI Driver
helm repo add aws-ebs-csi-driver https://kubernetes-sigs.github.io/aws-ebs-csi-driver
helm upgrade --install aws-ebs-csi-driver \
  aws-ebs-csi-driver/aws-ebs-csi-driver \
  --namespace kube-system \
  --set enableVolumeResizing=true \
  --set enableVolumeSnapshot=true

# GCP Filestore CSI Driver
kubectl apply -k "github.com/kubernetes-sigs/gcp-filestore-csi-driver/deploy/kubernetes/overlays/stable"

Longhorn: Open Source Distributed Storage__HTMLTAG_208___
# Kiểm tra prerequisites
curl -sfL https://raw.githubusercontent.com/longhorn/longhorn/v1.7.0/scripts/environment_check.sh | bash

# Cài Longhorn qua Helm
helm repo add longhorn https://charts.longhorn.io
helm repo update

helm install longhorn longhorn/longhorn \
  --namespace longhorn-system \
  --create-namespace \
  --version 1.7.0 \
  --set defaultSettings.defaultReplicaCount=3 \
  --set defaultSettings.storageMinimalAvailablePercentage=15 \
  --set defaultSettings.storageOverProvisioningPercentage=200

# Verify
kubectl get pods -n longhorn-system
kubectl get storageclass
# Longhorn StorageClass
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
  name: longhorn-fast
  annotations:
    storageclass.kubernetes.io/is-default-class: "false"
provisioner: driver.longhorn.io
parameters:
  numberOfReplicas: "3"
  staleReplicaTimeout: "30"
  fromBackup: ""
  fsType: ext4
  dataLocality: "best-effort"  # Prefer local replica
  diskSelector: "ssd"          # Hanya gunakan SSD disks
allowVolumeExpansion: true
reclaimPolicy: Delete
volumeBindingMode: WaitForFirstConsumer

Rook/Ceph: Enterprise Storage

# Cài Rook Ceph Operator
kubectl apply -f https://raw.githubusercontent.com/rook/rook/v1.15.0/deploy/examples/crds.yaml
kubectl apply -f https://raw.githubusercontent.com/rook/rook/v1.15.0/deploy/examples/common.yaml
kubectl apply -f https://raw.githubusercontent.com/rook/rook/v1.15.0/deploy/examples/operator.yaml

# Tạo Ceph cluster
kubectl apply -f https://raw.githubusercontent.com/rook/rook/v1.15.0/deploy/examples/cluster.yaml

# Verify
kubectl get CephCluster -n rook-ceph

Volume Snapshots

Concept__HTMLTAG_214___

Volume Snapshots is a GA feature from K8s 1.20 that allows creating point-in-time snapshots of PVCs. Requires CSI driver that supports snapshot capability and snapshot controller installed.

Install Snapshot Controller

# Cài snapshot CRDs và controller
kubectl apply -f https://raw.githubusercontent.com/kubernetes-csi/external-snapshotter/v8.0.0/client/config/crd/snapshot.storage.k8s.io_volumesnapshotclasses.yaml
kubectl apply -f https://raw.githubusercontent.com/kubernetes-csi/external-snapshotter/v8.0.0/client/config/crd/snapshot.storage.k8s.io_volumesnapshotcontents.yaml
kubectl apply -f https://raw.githubusercontent.com/kubernetes-csi/external-snapshotter/v8.0.0/client/config/crd/snapshot.storage.k8s.io_volumesnapshots.yaml

kubectl apply -f https://raw.githubusercontent.com/kubernetes-csi/external-snapshotter/v8.0.0/deploy/kubernetes/snapshot-controller/rbac-snapshot-controller.yaml
kubectl apply -f https://raw.githubusercontent.com/kubernetes-csi/external-snapshotter/v8.0.0/deploy/kubernetes/snapshot-controller/setup-snapshot-controller.yaml

VolumeSnapshotClass

apiVersion: snapshot.storage.k8s.io/v1
kind: VolumeSnapshotClass
metadata:
  name: csi-aws-vsc
  annotations:
    snapshot.storage.kubernetes.io/is-default-class: "true"
driver: ebs.csi.aws.com
deletionPolicy: Delete  # hoặc Retain
parameters:
  tagSpecification_1: "key=Name,value=k8s-snapshot"
  tagSpecification_2: "key=Environment,value=production"

Create VolumeSnapshot

apiVersion: snapshot.storage.k8s.io/v1
kind: VolumeSnapshot
metadata:
  name: postgres-snapshot-20260330
  namespace: production
  labels:
    backup-type: daily
    date: "2026-03-30"
spec:
  volumeSnapshotClassName: csi-aws-vsc
  source:
    persistentVolumeClaimName: database-storage  # PVC cần snapshot
# Xem trạng thái snapshot
kubectl get volumesnapshot -n production
kubectl describe volumesnapshot postgres-snapshot-20260330 -n production

# Output:
# Name:         postgres-snapshot-20260330
# Namespace:    production
# Status:
#   Bound Volume Snapshot Content Name: snapcontent-abc123
#   Creation Time: 2026-03-30T10:00:00Z
#   Ready To Use: true
#   Restore Size: 50Gi

Restore from Snapshot

# Tạo PVC mới từ snapshot
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: postgres-restored-20260330
  namespace: production
spec:
  accessModes:
  - ReadWriteOnce
  resources:
    requests:
      storage: 50Gi
  storageClassName: ebs-gp3
  dataSource:
    name: postgres-snapshot-20260330
    kind: VolumeSnapshot
    apiGroup: snapshot.storage.k8s.io
# Verify restored PVC
kubectl get pvc postgres-restored-20260330 -n production

# Deploy postgres với PVC được restored
kubectl patch deployment postgres -n production \
  --patch '{"spec":{"template":{"spec":{"volumes":[{"name":"postgres-storage","persistentVolumeClaim":{"claimName":"postgres-restored-20260330"}}]}}}}'

VolumeAttributesClass (K8s 1.29+)

The Problem With The Old Approach

Previously, if you wanted to change the IOPS or throughput of a volume (for example, from gp3 3000 IOPS to 10000 IOPS), you had to delete the PVC, create a new StorageClass, and create a new PVC — a complicated and downtime-inducing process.

VolumeAttributesClass (VAC) is a new API (Beta K8s 1.31) that allows changing mutable volume attributes such as IOPS and throughput without removing the PVC.

# Định nghĩa VolumeAttributesClass
apiVersion: storage.k8s.io/v1beta1
kind: VolumeAttributesClass
metadata:
  name: silver
driverName: ebs.csi.aws.com
parameters:
  type: gp3
  iops: "3000"
  throughput: "125"
---
apiVersion: storage.k8s.io/v1beta1
kind: VolumeAttributesClass
metadata:
  name: gold
driverName: ebs.csi.aws.com
parameters:
  type: gp3
  iops: "10000"
  throughput: "500"
# PVC ban đầu với silver class
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: database-storage
  namespace: production
spec:
  accessModes:
  - ReadWriteOnce
  resources:
    requests:
      storage: 100Gi
  storageClassName: ebs-gp3
  volumeAttributesClassName: silver  # Bắt đầu với silver
# Nâng cấp lên gold (tăng IOPS) không cần xóa PVC
kubectl patch pvc database-storage -n production \
  --patch '{"spec":{"volumeAttributesClassName":"gold"}}'

# Kiểm tra trạng thái modify
kubectl describe pvc database-storage -n production
# Events:
#   Normal ModifyVolumeComplete: Volume attributes successfully modified

Summary

Storage in Kubernetes has matured significantly with CSI being a mandatory platform since K8s 1.30-1.31. Points to remember:

  • emptyDir for ephemeral shared storage in Pod, PVC/PV for persistent data__HTMLTAG_245___
  • CSI required: no longer in-tree plugins from K8s 1.30+, must migrate to CSI drivers
  • StorageClass with dynamic provisioning is the standard way — admin defines the class, developer just needs to claim
  • Longhorn is a good choice for on-premise clusters with distributed replicated storage
  • Volume Snapshots allows point-in-time backup/restore, requires CSI driver and snapshot controller
  • VolumeAttributesClass (K8s 1.29+ Beta) allows changing IOPS/throughput without clearing PVC
  • Always use WaitForFirstConsumer binding mode with cloud volumes to avoid cross-AZ attachment issues