Kubernetes Storage Explained: PV, PVC, StorageClasses, StatefulSets
Kubernetes Storage Explained: PV, PVC, StorageClasses, StatefulSets
Containers are ephemeral — the filesystem dies with the pod. For anything that must outlive a restart (databases, uploads, queues), Kubernetes has a storage abstraction stack: Volumes, PersistentVolumes, PersistentVolumeClaims, and StorageClasses.
Table of Contents
- The pieces
- Dynamic provisioning: the normal way
- Access modes and reclaim policies
- StatefulSets: storage that follows the workload
- The mistakes that hurt
The pieces
- Volume — a directory mounted into a pod, lifecycle bound to the pod. emptyDir is the trivial case: scratch space that dies with the pod.
- PersistentVolume (PV) — a chunk of storage in the cluster, provisioned by an admin or dynamically.
- PersistentVolumeClaim (PVC) — a pod's request for storage: "give me 10 GiB, read-write, fast." Kubernetes binds the claim to a matching PV.
- StorageClass — the recipe for dynamic provisioning (EBS gp3, Azure Disk, Ceph, local-path-provisioner).
The split is deliberate: apps ask for storage (PVC), infrastructure supplies it (PV + StorageClass). App manifests stay portable across clouds.
Dynamic provisioning: the normal way
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: pgdata
spec:
accessModes: ["ReadWriteOnce"]
storageClassName: gp3
resources:
requests:
storage: 20Gi
spec:
containers:
- name: postgres
volumeMounts:
- { name: data, mountPath: /var/lib/postgresql/data }
volumes:
- name: data
persistentVolumeClaim:
claimName: pgdata
The PVC triggers the StorageClass's provisioner — disk created, PV bound, pod gets its mount. No administrator in the loop.
Access modes and reclaim policies
- ReadWriteOnce (RWO) — one node at a time. Most block storage (EBS, PD, Azure Disk).
- ReadWriteMany (RWX) — many nodes. Needs NFS/EFS/CephFS-style storage.
- ReadOnlyMany (ROX) — read-only, multi-node.
Reclaim policies on the PV: Delete (default for dynamic) destroys the disk when the PVC goes; Retain keeps it for manual recovery. For data you care about, Retain — or better, both a retention plan and tested backups.
StatefulSets: storage that follows the workload
Deployments + PVCs work for one replica. For replicated stateful systems, StatefulSets give each pod a stable identity (db-0, db-1), stable DNS, and its own PVC via volumeClaimTemplates:
apiVersion: apps/v1
kind: StatefulSet
metadata:
name: pg
spec:
serviceName: pg
replicas: 3
selector: { matchLabels: { app: pg } }
template:
metadata: { labels: { app: pg } }
spec:
containers:
- name: postgres
image: postgres:17
volumeMounts:
- { name: data, mountPath: /var/lib/postgresql/data }
volumeClaimTemplates:
- metadata: { name: data }
spec:
accessModes: ["ReadWriteOnce"]
storageClassName: gp3
resources: { requests: { storage: 20Gi } }
Pod pg-2 always gets claim data-pg-2. Reschedule it on any node and the same data reattaches.
The mistakes that hurt
- RWX on block storage — EBS can't multi-attach; your pods will pend forever.
- No storageClassName — inherits the cluster default, which may be slow or hostPath.
- Deleting a namespace thinking data survives — PVCs are namespaced; with Delete reclaim, the underlying disk goes too.
- Databases on whatever storage — match IOPS/latency class to workload; a chatty DB on network storage is misery.
Storage in Kubernetes rewards understanding the indirection once: pod → PVC → PV → StorageClass → actual disk. After that, stateful workloads stop being scary.
Related Articles
- Kubernetes for Beginners: Pods, Deployments, and Services
- Kubernetes Ingress and TLS: Exposing Services Properly
Last Updated: October 2026 Author: CloudOpsGuide Team Difficulty: Intermediate Estimated Reading Time: 12 minutes