Asterrr's Handbook

Pod placement

Steering Pods with nodeSelector, node affinity and Pod anti-affinity, keeping them off nodes with taints and tolerations, spreading them with topology spread constraints, and ordering them with PriorityClass and preemption.

Exam tasks: 2.5 (configure Pod admission and scheduling: node affinity and related placement controls)

The decision: must this Pod run on (or avoid) certain nodes, should it be attracted by labels or should other Pods be repelled by taints, and what wins when the cluster is full?

How the scheduler places a Pod

MechanismSet onEffectHard or soft
nodeNamePodSkips the scheduler entirely; the kubelet on that node runs itHard, ignores taints of effect NoSchedule
nodeSelectorPodNode must have all listed labelsHard
Node affinityPodLabel expressions with In, NotIn, Exists, DoesNotExist, Gt, Ltrequired... hard, preferred... weighted
Pod affinity / anti-affinityPodCo-locate with, or keep away from, Pods matching a selector, per topologyKeyBoth forms
TaintNodeRepels Pods that don't tolerate itNoSchedule, NoExecute hard; PreferNoSchedule soft
TolerationPodAllows (doesn't require) scheduling onto a tainted noden/a
Topology spreadPodCaps the imbalance of matching Pods across zones or nodesDoNotSchedule or ScheduleAnyway
PriorityClassPodOrder in the queue, and permission to preempt lower-priority Podsn/a

Labels, nodeSelector and node affinity

k label node worker-3 storage=nvme
k label node worker-3 storage-           # remove the label
k get nodes -L storage,topology.kubernetes.io/zone

Well-known labels you can rely on: kubernetes.io/hostname, kubernetes.io/os, kubernetes.io/arch, topology.kubernetes.io/zone, topology.kubernetes.io/region, node.kubernetes.io/instance-type.

spec:
  nodeSelector:
    kubernetes.io/os: linux
  affinity:
    nodeAffinity:
      requiredDuringSchedulingIgnoredDuringExecution:
        nodeSelectorTerms:            # terms are ORed
        - matchExpressions:           # expressions inside one term are ANDed
          - key: storage
            operator: In
            values: ["nvme", "ssd"]
      preferredDuringSchedulingIgnoredDuringExecution:
      - weight: 50                    # 1 to 100
        preference:
          matchExpressions:
          - key: topology.kubernetes.io/zone
            operator: In
            values: ["zone-b"]
  • IgnoredDuringExecution: removing the label later doesn't evict running Pods.
  • nodeSelector and node affinity together must both match.
  • NotIn and DoesNotExist express "keep off these nodes" without touching the nodes.

Exam signal

"Must run only on nodes labelled X" is nodeSelector or required node affinity. "Should prefer" is preferredDuringScheduling.... "Only these Pods may use the GPU nodes" needs a taint on the nodes plus a toleration and an affinity on the Pods.

Pod affinity and anti-affinity

affinity:
  podAntiAffinity:
    requiredDuringSchedulingIgnoredDuringExecution:
    - labelSelector:
        matchLabels:
          app: sentry-cache
      topologyKey: kubernetes.io/hostname    # at most one sentry-cache Pod per node
  podAffinity:
    preferredDuringSchedulingIgnoredDuringExecution:
    - weight: 80
      podAffinityTerm:
        labelSelector:
          matchLabels:
            app: sentry-api
        topologyKey: topology.kubernetes.io/zone   # same zone as the API if possible
  • topologyKey defines the "place": a node (hostname) or a zone. It must be a label on the nodes.
  • Required anti-affinity per node with more replicas than nodes leaves the extras Pending. Use the preferred form, or topology spread, when replicas can outnumber nodes.

Taints and tolerations

k taint nodes worker-4 dedicated=analytics:NoSchedule
k taint nodes worker-4 dedicated=analytics:NoSchedule-     # remove (note the trailing -)
k describe node worker-4 | grep -i taints
tolerations:
- key: dedicated
  operator: Equal          # Exists matches any value; omit value with Exists
  value: analytics
  effect: NoSchedule
- key: node.kubernetes.io/unreachable
  operator: Exists
  effect: NoExecute
  tolerationSeconds: 60    # leave an unreachable node after 60 s instead of 300 s
EffectNew Pods without tolerationRunning Pods without toleration
PreferNoScheduleAvoided if possibleStay
NoScheduleNot scheduledStay
NoExecuteNot scheduledEvicted (after tolerationSeconds, if set in their toleration)
  • kubeadm taints control plane nodes with node-role.kubernetes.io/control-plane:NoSchedule. DaemonSets that must run there need that toleration.
  • The node controller adds node.kubernetes.io/not-ready and node.kubernetes.io/unreachable NoExecute taints to failing nodes. Every Pod gets default tolerations for them of 300 seconds, which is why Pods take about 5 minutes to move off a dead node.

A toleration as a node selector

A toleration only permits a Pod onto a tainted node; the scheduler is still free to put it anywhere else. For "these Pods must run on the analytics nodes and nothing else may", you need both halves: the taint keeps others off, and node affinity (or nodeSelector) pins your Pods on.

Topology spread constraints

topologySpreadConstraints:
- maxSkew: 1
  topologyKey: topology.kubernetes.io/zone
  whenUnsatisfiable: DoNotSchedule   # or ScheduleAnyway (soft)
  labelSelector:
    matchLabels:
      app: sentry-api
  matchLabelKeys: ["pod-template-hash"]   # count only Pods of the same rollout revision
  • maxSkew is the largest allowed difference in matching Pods between the fullest and emptiest domain.
  • Unlike required anti-affinity, it scales past one Pod per domain: 9 replicas in 3 zones go 3/3/3.
  • minDomains (with DoNotSchedule) treats missing domains as having zero Pods, forcing spread across at least that many.

PriorityClass and preemption

k create priorityclass payments-critical --value=100000 --description="customer-facing payments"
k create priorityclass batch-low --value=100 --preemption-policy=Never
spec:
  priorityClassName: payments-critical
  • Higher value schedules first. When no node fits, the scheduler may evict lower-priority Pods to make room (preemption). PDBs are respected on a best-effort basis only.
  • preemptionPolicy: Never keeps the queue priority but never evicts others.
  • globalDefault: true on one class sets the priority for Pods without a class; otherwise they get 0.
  • Built-in system-cluster-critical and system-node-critical (about 2 billion) are reserved for cluster components.
300 s
Default toleration for not-ready and unreachable taints before Pods are evicted.
1–100
Weight range for preferred affinity terms.
OR / AND
nodeSelectorTerms are ORed; matchExpressions within a term are ANDed.
0
Priority of a Pod with no class when no globalDefault exists.

Scenarios

Scenario
Nodes gpu-1 and gpu-2 must run only Pods of Deployment `trainer`, and `trainer` Pods must run only on those nodes. Both nodes carry the label accel=gpu. Which combination achieves this?
Scenario
A 3-node cluster runs Deployment `sentinel` with required podAntiAffinity on topologyKey kubernetes.io/hostname. After scaling it from 3 to 5 replicas, two Pods stay Pending with 'didn't match pod anti-affinity rules'. The team wants all 5 running but as evenly spread as possible. What should you change?
Scenario
A node was tainted with `maintenance=true:NoSchedule`. An hour later its existing Pods are all still running on it, and the admin expected them to have moved. Why, and what would move them?

Drill

kubectl config use-context drill-w6. In namespace harbor:

  1. Label node worker-2 with tier=edge and taint it tier=edge:NoSchedule.
  2. Create Deployment beacon (image nginx:1.27-alpine, 2 replicas) whose Pods run only on tier=edge nodes.
  3. Create PriorityClass beacon-high with value 50000 and use it for beacon.

Further reading

On this page