Asterrr's Handbook

Kubernetes storage

Ephemeral and persistent volumes, PersistentVolumes and PersistentVolumeClaims, access modes and reclaim policies, StorageClasses and dynamic provisioning, CSI drivers, and StatefulSet storage for KCNA.

Exam tasks: 2.4 (storage)

The decision: does this data need to outlive the container, the Pod or the node, and which object provides storage with that lifetime?

Lifetimes first

Container filesystems are ephemeral: when a container restarts, anything it wrote to its own filesystem is gone. Volumes fix that, with different lifetimes.

Gone with the PodOutlives everything
  1. Container filesystem
    container
    Lost on every container restart.
  2. emptyDir
    Pod
    Survives container restarts, shared by containers in the Pod, deleted with the Pod.
  3. hostPath
    node
    A directory on the node. Stays on that node only, and is a security risk.
  4. PersistentVolume
    independent
    Network or cloud storage with its own lifecycle, follows the Pod to any node that can attach it.
Volume typeUse for
emptyDirScratch space, caches, sharing files between a main container and a sidecar. medium: Memory makes it a RAM-backed tmpfs
configMap, secretConfiguration files and credentials projected as files
projectedSeveral sources (ServiceAccount token, ConfigMap, Secret, downward API) in one directory
hostPathNode agents that must read host files (log collectors). Blocked by the Baseline Pod Security profile
persistentVolumeClaimAny data that must survive Pod deletion: databases, uploads, queues

PV and PVC: the split

  • A PersistentVolume (PV) is a piece of storage in the cluster. It's cluster-scoped.
  • A PersistentVolumeClaim (PVC) is a request for storage (size, access mode, class). It's namespaced, and Pods refer to it by name.
  • The control plane binds each PVC to exactly one PV that satisfies it. A PVC with no match stays Pending.
  • The split separates concerns: app teams ask for "20Gi of fast disk" without knowing which backend provides it.

Access modes

ModeShortMeaning
ReadWriteOnceRWORead-write by a single node (several Pods on that node can share it)
ReadOnlyManyROXRead-only by many nodes
ReadWriteManyRWXRead-write by many nodes. Needs shared file storage such as NFS or CephFS
ReadWriteOncePodRWOPRead-write by a single Pod in the whole cluster. CSI volumes only
  • A volume supports only the modes its backend allows. Most block storage (cloud disks) is RWO only.

RWO means one Pod

ReadWriteOnce restricts the volume to one node, not one Pod. Two Pods scheduled on the same node can both mount it. If the requirement is "only one Pod may ever write", the answer is ReadWriteOncePod.

Reclaim policy

What happens to the PV and its data when the PVC is deleted:

  • Delete: the PV and the underlying disk are deleted. The default for dynamically provisioned volumes.
  • Retain: the PV stays (status Released) with its data. An admin must clean up or reuse it manually.

Legacy: use dynamic provisioning with the Delete or Retain policy instead

The Recycle reclaim policy (a basic rm -rf on the volume) is deprecated. Use dynamic provisioning instead.

StorageClasses and dynamic provisioning

Static provisioning: an admin creates PVs in advance and PVCs bind to them. Dynamic provisioning: the PVC names a StorageClass, and its provisioner creates a matching PV on demand. Dynamic is the norm.

apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
  name: fast-ssd
provisioner: ebs.csi.aws.com
parameters:
  type: gp3
reclaimPolicy: Delete
volumeBindingMode: WaitForFirstConsumer
allowVolumeExpansion: true
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: orders-db-data
  namespace: shop
spec:
  storageClassName: fast-ssd
  accessModes: ["ReadWriteOnce"]
  resources:
    requests:
      storage: 20Gi
  • One StorageClass can be marked default (annotation storageclass.kubernetes.io/is-default-class: "true"). PVCs that don't name a class get it.
  • volumeBindingMode: WaitForFirstConsumer delays creating the disk until a Pod using the PVC is scheduled, so the disk lands in the same zone as the Pod. Immediate creates it right away and can strand it in the wrong zone.
  • allowVolumeExpansion: true lets you grow a PVC by editing its size. Shrinking isn't supported.
  • VolumeSnapshots (with a VolumeSnapshotClass) take point-in-time copies through CSI drivers that support it.

CSI: the storage plug-in standard

  • The Container Storage Interface is a standard API between container orchestrators and storage systems. Vendors ship a CSI driver (controller plus a node DaemonSet) instead of putting code into Kubernetes.
  • Drivers handle create, delete, attach, mount, snapshot and resize for their backend.
  • The older in-tree volume plugins for cloud disks have been migrated to CSI and removed from Kubernetes.
  • CSI sits alongside the other plug-in interfaces: CRI for runtimes and CNI for networking.

Exam signal

"Let a storage vendor support Kubernetes without changing Kubernetes code" is CSI. "Create volumes automatically when a developer asks for one" is dynamic provisioning via a StorageClass. "Request storage without knowing the backend" is a PVC.

StatefulSet storage

  • A StatefulSet's volumeClaimTemplates creates one PVC per replica, named <template>-<statefulset>-<ordinal>, for example data-ledger-0, data-ledger-1.
  • When a replica is rescheduled, it gets the same PVC back, so each replica keeps its own data.
  • By default, PVCs are kept when you scale down or delete the StatefulSet, to protect data. persistentVolumeClaimRetentionPolicy can change that to delete them.
  • A Deployment with one PVC shares the same claim across all replicas, which only works with RWX storage.
Cluster-scoped
PersistentVolume and StorageClass.
Namespaced
PersistentVolumeClaim.
4 access modes
RWO, ROX, RWX, RWOP.
Delete
Default reclaim policy for dynamically provisioned PVs.
1 PVC per replica
From a StatefulSet's volumeClaimTemplates.

For hands-on depth, see the CKA pages Persistent volumes and Storage classes.

Scenarios

Scenario
A developer needs 50Gi of disk for a database Pod and does not know which storage backend the cluster uses. The cluster has a default StorageClass. What should the developer create?
Scenario
A cluster spans three availability zones. Dynamically provisioned volumes are sometimes created in a zone where the Pod cannot be scheduled, leaving the Pod Pending. Which StorageClass setting fixes this?
Scenario
A three-replica message broker must give each replica its own persistent disk, and a replica must get its own disk back after being rescheduled. Which approach fits?

Further reading

On this page