Asterrr's Handbook

StorageClasses and dynamic provisioning

How a StorageClass and its CSI driver create PVs on demand, setting the default class, Immediate vs WaitForFirstConsumer binding, allowVolumeExpansion and growing a claim, and which fields you can't change later.

Exam tasks: 4.1 (StorageClasses and dynamic provisioning)

The decision: which class does a claim end up with, when is the real volume created, and will you be able to grow it later?

What happens when a claim names a class

  • The provisioner named in the class does the work. For CSI drivers it's the external-provisioner sidecar running next to the driver's controller; it watches claims whose class names its driver.
  • The new PV copies reclaimPolicy, mountOptions and the class name from the StorageClass, and is pre-bound to the claim that triggered it.
  • kubectl get csidrivers lists installed CSI drivers. kubectl get sc shows each class's provisioner, reclaim policy, binding mode and whether expansion is allowed.

Anatomy of a StorageClass

apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
  name: fast-ssd
  annotations:
    storageclass.kubernetes.io/is-default-class: "true"   # optional
provisioner: csi.quarrystore.example      # driver name, required
parameters:                               # opaque to Kubernetes, driver-specific
  tier: nvme
  replicas: "2"
reclaimPolicy: Retain                     # Delete if omitted
volumeBindingMode: WaitForFirstConsumer   # Immediate if omitted
allowVolumeExpansion: true                # false if omitted
mountOptions:
  - noatime
FieldDefaultChange later?Notes
provisionernone, requiredNoMust match a running driver, or claims stay Pending with no error from the API
parametersemptyNoValues are strings; quote numbers. Unknown keys are the driver's problem, not the API's
reclaimPolicyDeleteNoApplies only to PVs created after you set it
volumeBindingModeImmediateNoSee below
allowVolumeExpansionfalseYesTurning it on lets existing claims of this class grow
mountOptionsnoneYesNot validated; a bad option fails at mount time

Editing the class to fix existing volumes

A class is a template. Changing or recreating it never touches PVs it already made: their reclaim policy and parameters were copied at creation. To keep data from an existing dynamic volume, patch the PV's persistentVolumeReclaimPolicy to Retain. To change a fixed class field, delete the class and recreate it with the same name; existing claims and PVs keep working.

The default class

  • Mark a class as default with the annotation storageclass.kubernetes.io/is-default-class: "true". Any claim created without a storageClassName gets that class written into it at admission.
  • If more than one class carries the annotation, new claims get the most recently created default. Keep exactly one; multiple defaults exist only to make migrations painless.
  • Claims created while no default existed stay unset, and are updated retroactively once a default appears.
  • A claim with storageClassName: "" opts out: it is never defaulted and only binds classless PVs.
# Move the default from local-path to fast-ssd
k patch sc local-path -p '{"metadata":{"annotations":{"storageclass.kubernetes.io/is-default-class":"false"}}}'
k patch sc fast-ssd   -p '{"metadata":{"annotations":{"storageclass.kubernetes.io/is-default-class":"true"}}}'
k get sc        # "(default)" appears after the name

Exam signal

"Make X the default StorageClass" is two edits, not one: set the annotation to "true" on X and remove it or set it to "false" on the old default. The grader checks that only one class shows (default). The value must be the quoted string "true".

Binding mode

ImmediateWaitForFirstConsumer
Volume createdAs soon as the PVC existsWhen a Pod using the PVC is scheduled
PVC status before a Pod existsBoundPending, event WaitForFirstConsumer. This is normal
TopologyChosen without knowing the Pod, so the disk may land in a zone the Pod can't useFollows the node the scheduler picked, and respects the Pod's affinity and resource needs
Use forNetwork storage reachable from every nodeZonal disks, local volumes, anything node- or zone-bound
  • For local PVs there's no dynamic provisioning: use provisioner: kubernetes.io/no-provisioner with WaitForFirstConsumer, so binding waits until the scheduler knows which node's disk the Pod can use.
  • allowedTopologies restricts where volumes may be provisioned (for example, two zones only). Usually you leave it out and let WaitForFirstConsumer follow the Pod.

nodeName with WaitForFirstConsumer

Setting spec.nodeName on a Pod skips the scheduler, so nothing annotates the claim with a selected node and a WaitForFirstConsumer claim never binds. Steer the Pod with nodeSelector or node affinity instead.

Growing a volume

  • Only claims bound to an expandable volume type (CSI drivers that support it) can grow. The driver must implement expansion; the field alone doesn't make it possible.
  • Never shrink. A request smaller than the current size is rejected. You can't edit the class of a bound claim either: copy data to a new claim instead.
  • Watch kubectl describe pvc for conditions such as Resizing and FileSystemResizePending. File system growth finishes when a Pod uses the claim, so an unused claim waits.
  • If an expansion fails because the backend has no room, you may retry with a smaller value than you asked for, as long as it's still above the current status.capacity.

Editing the PV instead of the PVC

Raising spec.capacity on the PV and then matching it on the claim makes Kubernetes think the resize already happened, and nothing expands the real disk. Always change the claim's request.

Delete
Default reclaimPolicy of a StorageClass.
Immediate
Default volumeBindingMode.
false
Default allowVolumeExpansion. Expansion is grow-only.
Newest wins
With several default classes, new claims get the most recently created one.
storage.k8s.io/v1
API group and version for StorageClass and CSIDriver objects.

Legacy: use CSI driver names and the GA annotation instead

In-tree provisioners such as kubernetes.io/aws-ebs and kubernetes.io/gce-pd are redirected to CSI drivers (ebs.csi.aws.com, pd.csi.storage.gke.io) through CSI migration; write new classes against the CSI name. The annotation storageclass.beta.kubernetes.io/is-default-class and the PVC annotation volume.beta.kubernetes.io/storage-class are replaced by storageclass.kubernetes.io/is-default-class and the storageClassName field.

Scenarios

Scenario
A cluster spans three zones. Class 'zonal-block' uses a CSI driver for zonal disks with default settings. A StatefulSet whose Pods require nodes labelled gpu=true (present in one zone only) keeps a replica Pending with a volume node affinity conflict. What is the BEST fix?
Scenario
A developer reports that PVC 'quill-cache' in namespace scribe has been Pending for ten minutes. Its class 'node-scratch' uses provisioner rancher.io/local-path and WaitForFirstConsumer. `kubectl get pods -n scribe` shows no Pods using the claim. What should you tell them?
Scenario
PVC 'atlas-db' (class 'standard-rwo', 10Gi, Bound, mounted by a running Pod) needs 25Gi. The edit to 25Gi is rejected with a message that the claim's class does not support resize. The CSI driver supports online expansion. What is the least disruptive fix?

Drill

kubectl config use-context cka-provision

The cluster runs the rancher.io/local-path provisioner, and class local-path is currently the default.

  1. Create a StorageClass scratch-wffc with that provisioner, reclaim policy Retain, binding mode WaitForFirstConsumer, and volume expansion allowed.
  2. Make scratch-wffc the only default StorageClass.
  3. In namespace kiln, create PVC glaze-data requesting 256Mi ReadWriteOnce without naming a class, and confirm it got scratch-wffc.
  4. Start a Pod glaze (image nginx:1.27) that mounts it at /usr/share/nginx/html, and confirm the claim binds.

Further reading

On this page