Persistent Storage and CSI in Kubernetes__HTMLTAG_66___
Containers are ephemeral — when a container restarts or Pod is rescheduled on another node, all data inside is lost. This is a serious problem with stateful workloads like databases, message queues, and file storage. Kubernetes solves this problem with a rich storage system, and from Kubernetes 1.30-1.31, in-tree storage plugins have been completely removed, replaced by Container Storage Interface (CSI) standardized drivers.
Storage Lifecycle: From Ephemeral to Persistent__HTMLTAG_72___
emptyDir: Ephemeral Storage
emptyDir creates an empty directory when the Pod is created and persists for the life of the Pod. When the Pod is deleted, the data is completely lost. Suitable for caches, temporary files, sharing data between containers in the same Pod.
apiVersion: v1
kind: Pod
metadata:
name: shared-storage-pod
spec:
containers:
- name: writer
image: busybox
command: ["/bin/sh", "-c"]
args: ["while true; do date >> /shared/log.txt; sleep 1; done"]
volumeMounts:
- name: shared-data
mountPath: /shared
- name: reader
image: busybox
command: ["/bin/sh", "-c"]
args: ["tail -f /shared/log.txt"]
volumeMounts:
- name: shared-data
mountPath: /shared
volumes:
- name: shared-data
emptyDir:
sizeLimit: 500Mi # Giới hạn kích thước
medium: Memory # Lưu trong RAM (tmpfs) - nhanh hơn nhưng dùng memory
hostPath: Node-Local Storage
hostPath mounts a path on the host node to the container. The data exists when the container restarts but is lost when the Pod is scheduled to another node. Should not be used in production unless there is a special reason (system daemons, monitoring agents).
apiVersion: v1
kind: Pod
metadata:
name: host-path-pod
spec:
containers:
- name: app
image: nginx:1.27
volumeMounts:
- name: host-volume
mountPath: /data
volumes:
- name: host-volume
hostPath:
path: /data/myapp
type: DirectoryOrCreate # Tạo directory nếu chưa tồn tại
PersistentVolume and PersistentVolumeClaim
Core Concepts
Kubernetes separates providing storage (admin) and using storage (developer) through two abstractions:
- PersistentVolume (PV): Cluster-level storage resource created by admin or dynamically provisioned. Represents a piece of actual storage (NFS share, cloud disk, etc.)
- PersistentVolumeClaim (PVC): User request for storage. Developers just need to declare "I need 10Gi of storage with ReadWriteOnce" without knowing where the storage is.
PersistentVolume Definition
apiVersion: v1
kind: PersistentVolume
metadata:
name: pv-database-001
labels:
type: ssd
environment: production
spec:
capacity:
storage: 100Gi
volumeMode: Filesystem # hoặc Block
accessModes:
- ReadWriteOnce # Chỉ một node mount được tại một thời điểm
persistentVolumeReclaimPolicy: Retain # Giữ lại data sau khi PVC released
storageClassName: premium-ssd
# CSI volume source (thay vì in-tree)
csi:
driver: ebs.csi.aws.com
volumeHandle: vol-0a1b2c3d4e5f67890
fsType: ext4
volumeAttributes:
storage.kubernetes.io/csiProvisionerIdentity: "1234567890"
PersistentVolumeClaim
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: database-storage
namespace: production
spec:
accessModes:
- ReadWriteOnce
volumeMode: Filesystem
resources:
requests:
storage: 50Gi
storageClassName: premium-ssd
# Selector để bind với specific PV (optional)
selector:
matchLabels:
type: ssd
environment: production
# Xem trạng thái PVC
kubectl get pvc -n production
# Output:
# NAME STATUS VOLUME CAPACITY ACCESS MODES STORAGECLASS AGE
# database-storage Bound pv-database-001 100Gi RWO premium-ssd 5m
Using PVC in Pod
apiVersion: apps/v1
kind: Deployment
metadata:
name: postgres
namespace: production
spec:
replicas: 1
selector:
matchLabels:
app: postgres
template:
metadata:
labels:
app: postgres
spec:
containers:
- name: postgres
image: postgres:16
env:
- name: POSTGRES_PASSWORD
valueFrom:
secretKeyRef:
name: postgres-secret
key: password
- name: PGDATA
value: /var/lib/postgresql/data/pgdata
volumeMounts:
- name: postgres-storage
mountPath: /var/lib/postgresql/data
volumes:
- name: postgres-storage
persistentVolumeClaim:
claimName: database-storage
Access Modes
Kubernetes defines 4 access modes, showing how volumes can be mounted:
- ReadWriteOnce (RWO): A node that can mount read-write. Most common with block storage (EBS, GCE PD). From K8s 1.22+, RWO allows multiple Pods on the SAME read-write node.
- ReadOnlyMany (ROX): Multiple nodes can mount read-only at the same time. Matches shared config/data.
- ReadWriteMany (RWX): Many nodes mount read-write. Requires network filesystem such as NFS, CephFS, Azure Files. Important: block storage (EBS, GCE PD) does NOT support RWX.
- ReadWriteOncePod (RWOP): Only a single Pod in the entire cluster can mount. Stronger than RWO — ensures exclusive access at Pod level, not just node level. Requires CSI driver support.
StorageClass: Dynamic Provisioning
Why StorageClass?
Instead of admins having to create PVs manually, StorageClass allows dynamic provisioning — Kubernetes automatically creates PVs when there is a PVC request. Admin only needs to define storage "class" (eg: ssd-fast, hdd-cheap, nfs-shared).
# StorageClass với AWS EBS CSI Driver
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
name: ebs-gp3
annotations:
storageclass.kubernetes.io/is-default-class: "true"
provisioner: ebs.csi.aws.com
parameters:
type: gp3
iops: "3000"
throughput: "125"
encrypted: "true"
kmsKeyId: arn:aws:kms:ap-southeast-1:123456789012:key/mrk-abc123
volumeBindingMode: WaitForFirstConsumer # Đợi Pod được schedule trước khi tạo volume
reclaimPolicy: Delete # Xóa PV khi PVC bị xóa
allowVolumeExpansion: true # Cho phép resize PVC
---
# StorageClass cho Longhorn
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
name: longhorn-replicated
provisioner: driver.longhorn.io
parameters:
numberOfReplicas: "3"
staleReplicaTimeout: "2880"
fromBackup: ""
fsType: "ext4"
volumeBindingMode: Immediate
reclaimPolicy: Delete
allowVolumeExpansion: true
Reclaim Policy
- Delete: PV and underlying storage are deleted when PVC is deleted. Default with dynamic provisioning. Suitable for ephemeral workloads.
- Retain: The PV is retained (Released state) when the PVC is deleted. Admin must reclaim manually. Suitable for production databases that need data protection.
- Recycle: Deprecated, not recommended.
CSI: Container Storage Interface
Why Was In-Tree Plugins Removed?
Previously, Kubernetes had many in-tree storage plugins compiled directly into the core Kubernetes binary (aws-ebs, gce-pd, azure-disk, cephfs, nfs...). This creates many problems:
- Bug in storage plugin can crash entire kube-apiserver/kubelet
- Release cycle of storage drivers tied to Kubernetes release
- Difficult to maintain when the number of plugins increases
- Storage vendors cannot ship fixes independently
CSI (Container Storage Interface) solves all of these problems by standardizing the interface between Kubernetes and storage providers. Drivers run as separate Pods, can be updated independently.
Timeline Remove In-Tree Plugins__HTMLTAG_180___
- K8s 1.26-1.28: Many in-tree plugins deprecated
- K8s 1.29: In-tree NFS and many plugins converted to deprecated, CSI required
- K8s 1.30: In-tree NFS plugin removed
- K8s 1.31: In-tree CephFS, Ceph RBD plugins completely removed__HTMLTAG_197___
- K8s 1.32+: Continue removing the remaining in-tree plugins
Popular CSI Drivers
Cloud Providers
# AWS EBS CSI Driver
helm repo add aws-ebs-csi-driver https://kubernetes-sigs.github.io/aws-ebs-csi-driver
helm upgrade --install aws-ebs-csi-driver \
aws-ebs-csi-driver/aws-ebs-csi-driver \
--namespace kube-system \
--set enableVolumeResizing=true \
--set enableVolumeSnapshot=true
# GCP Filestore CSI Driver
kubectl apply -k "github.com/kubernetes-sigs/gcp-filestore-csi-driver/deploy/kubernetes/overlays/stable"
Longhorn: Open Source Distributed Storage__HTMLTAG_208___
# Kiểm tra prerequisites
curl -sfL https://raw.githubusercontent.com/longhorn/longhorn/v1.7.0/scripts/environment_check.sh | bash
# Cài Longhorn qua Helm
helm repo add longhorn https://charts.longhorn.io
helm repo update
helm install longhorn longhorn/longhorn \
--namespace longhorn-system \
--create-namespace \
--version 1.7.0 \
--set defaultSettings.defaultReplicaCount=3 \
--set defaultSettings.storageMinimalAvailablePercentage=15 \
--set defaultSettings.storageOverProvisioningPercentage=200
# Verify
kubectl get pods -n longhorn-system
kubectl get storageclass
# Longhorn StorageClass
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
name: longhorn-fast
annotations:
storageclass.kubernetes.io/is-default-class: "false"
provisioner: driver.longhorn.io
parameters:
numberOfReplicas: "3"
staleReplicaTimeout: "30"
fromBackup: ""
fsType: ext4
dataLocality: "best-effort" # Prefer local replica
diskSelector: "ssd" # Hanya gunakan SSD disks
allowVolumeExpansion: true
reclaimPolicy: Delete
volumeBindingMode: WaitForFirstConsumer
# Kiểm tra prerequisites
curl -sfL https://raw.githubusercontent.com/longhorn/longhorn/v1.7.0/scripts/environment_check.sh | bash
# Cài Longhorn qua Helm
helm repo add longhorn https://charts.longhorn.io
helm repo update
helm install longhorn longhorn/longhorn \
--namespace longhorn-system \
--create-namespace \
--version 1.7.0 \
--set defaultSettings.defaultReplicaCount=3 \
--set defaultSettings.storageMinimalAvailablePercentage=15 \
--set defaultSettings.storageOverProvisioningPercentage=200
# Verify
kubectl get pods -n longhorn-system
kubectl get storageclass# Longhorn StorageClass
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
name: longhorn-fast
annotations:
storageclass.kubernetes.io/is-default-class: "false"
provisioner: driver.longhorn.io
parameters:
numberOfReplicas: "3"
staleReplicaTimeout: "30"
fromBackup: ""
fsType: ext4
dataLocality: "best-effort" # Prefer local replica
diskSelector: "ssd" # Hanya gunakan SSD disks
allowVolumeExpansion: true
reclaimPolicy: Delete
volumeBindingMode: WaitForFirstConsumerRook/Ceph: Enterprise Storage
# Cài Rook Ceph Operator
kubectl apply -f https://raw.githubusercontent.com/rook/rook/v1.15.0/deploy/examples/crds.yaml
kubectl apply -f https://raw.githubusercontent.com/rook/rook/v1.15.0/deploy/examples/common.yaml
kubectl apply -f https://raw.githubusercontent.com/rook/rook/v1.15.0/deploy/examples/operator.yaml
# Tạo Ceph cluster
kubectl apply -f https://raw.githubusercontent.com/rook/rook/v1.15.0/deploy/examples/cluster.yaml
# Verify
kubectl get CephCluster -n rook-ceph
Volume Snapshots
Concept__HTMLTAG_214___
Volume Snapshots is a GA feature from K8s 1.20 that allows creating point-in-time snapshots of PVCs. Requires CSI driver that supports snapshot capability and snapshot controller installed.
Install Snapshot Controller
# Cài snapshot CRDs và controller
kubectl apply -f https://raw.githubusercontent.com/kubernetes-csi/external-snapshotter/v8.0.0/client/config/crd/snapshot.storage.k8s.io_volumesnapshotclasses.yaml
kubectl apply -f https://raw.githubusercontent.com/kubernetes-csi/external-snapshotter/v8.0.0/client/config/crd/snapshot.storage.k8s.io_volumesnapshotcontents.yaml
kubectl apply -f https://raw.githubusercontent.com/kubernetes-csi/external-snapshotter/v8.0.0/client/config/crd/snapshot.storage.k8s.io_volumesnapshots.yaml
kubectl apply -f https://raw.githubusercontent.com/kubernetes-csi/external-snapshotter/v8.0.0/deploy/kubernetes/snapshot-controller/rbac-snapshot-controller.yaml
kubectl apply -f https://raw.githubusercontent.com/kubernetes-csi/external-snapshotter/v8.0.0/deploy/kubernetes/snapshot-controller/setup-snapshot-controller.yaml
VolumeSnapshotClass
apiVersion: snapshot.storage.k8s.io/v1
kind: VolumeSnapshotClass
metadata:
name: csi-aws-vsc
annotations:
snapshot.storage.kubernetes.io/is-default-class: "true"
driver: ebs.csi.aws.com
deletionPolicy: Delete # hoặc Retain
parameters:
tagSpecification_1: "key=Name,value=k8s-snapshot"
tagSpecification_2: "key=Environment,value=production"
Create VolumeSnapshot
apiVersion: snapshot.storage.k8s.io/v1
kind: VolumeSnapshot
metadata:
name: postgres-snapshot-20260330
namespace: production
labels:
backup-type: daily
date: "2026-03-30"
spec:
volumeSnapshotClassName: csi-aws-vsc
source:
persistentVolumeClaimName: database-storage # PVC cần snapshot
# Xem trạng thái snapshot
kubectl get volumesnapshot -n production
kubectl describe volumesnapshot postgres-snapshot-20260330 -n production
# Output:
# Name: postgres-snapshot-20260330
# Namespace: production
# Status:
# Bound Volume Snapshot Content Name: snapcontent-abc123
# Creation Time: 2026-03-30T10:00:00Z
# Ready To Use: true
# Restore Size: 50Gi
Restore from Snapshot
# Tạo PVC mới từ snapshot
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: postgres-restored-20260330
namespace: production
spec:
accessModes:
- ReadWriteOnce
resources:
requests:
storage: 50Gi
storageClassName: ebs-gp3
dataSource:
name: postgres-snapshot-20260330
kind: VolumeSnapshot
apiGroup: snapshot.storage.k8s.io
# Verify restored PVC
kubectl get pvc postgres-restored-20260330 -n production
# Deploy postgres với PVC được restored
kubectl patch deployment postgres -n production \
--patch '{"spec":{"template":{"spec":{"volumes":[{"name":"postgres-storage","persistentVolumeClaim":{"claimName":"postgres-restored-20260330"}}]}}}}'
VolumeAttributesClass (K8s 1.29+)
The Problem With The Old Approach
Previously, if you wanted to change the IOPS or throughput of a volume (for example, from gp3 3000 IOPS to 10000 IOPS), you had to delete the PVC, create a new StorageClass, and create a new PVC — a complicated and downtime-inducing process.
VolumeAttributesClass (VAC) is a new API (Beta K8s 1.31) that allows changing mutable volume attributes such as IOPS and throughput without removing the PVC.
# Định nghĩa VolumeAttributesClass
apiVersion: storage.k8s.io/v1beta1
kind: VolumeAttributesClass
metadata:
name: silver
driverName: ebs.csi.aws.com
parameters:
type: gp3
iops: "3000"
throughput: "125"
---
apiVersion: storage.k8s.io/v1beta1
kind: VolumeAttributesClass
metadata:
name: gold
driverName: ebs.csi.aws.com
parameters:
type: gp3
iops: "10000"
throughput: "500"
# PVC ban đầu với silver class
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: database-storage
namespace: production
spec:
accessModes:
- ReadWriteOnce
resources:
requests:
storage: 100Gi
storageClassName: ebs-gp3
volumeAttributesClassName: silver # Bắt đầu với silver
# Nâng cấp lên gold (tăng IOPS) không cần xóa PVC
kubectl patch pvc database-storage -n production \
--patch '{"spec":{"volumeAttributesClassName":"gold"}}'
# Kiểm tra trạng thái modify
kubectl describe pvc database-storage -n production
# Events:
# Normal ModifyVolumeComplete: Volume attributes successfully modified
Summary
Storage in Kubernetes has matured significantly with CSI being a mandatory platform since K8s 1.30-1.31. Points to remember:
- emptyDir for ephemeral shared storage in Pod, PVC/PV for persistent data__HTMLTAG_245___
- CSI required: no longer in-tree plugins from K8s 1.30+, must migrate to CSI drivers
- StorageClass with dynamic provisioning is the standard way — admin defines the class, developer just needs to claim
- Longhorn is a good choice for on-premise clusters with distributed replicated storage
- Volume Snapshots allows point-in-time backup/restore, requires CSI driver and snapshot controller
- VolumeAttributesClass (K8s 1.29+ Beta) allows changing IOPS/throughput without clearing PVC
- Always use
WaitForFirstConsumerbinding mode with cloud volumes to avoid cross-AZ attachment issues