Pod placement
Steering Pods with nodeSelector, node affinity and Pod anti-affinity, keeping them off nodes with taints and tolerations, spreading them with topology spread constraints, and ordering them with PriorityClass and preemption.
Exam tasks: 2.5 (configure Pod admission and scheduling: node affinity and related placement controls)
The decision: must this Pod run on (or avoid) certain nodes, should it be attracted by labels or should other Pods be repelled by taints, and what wins when the cluster is full?
How the scheduler places a Pod
| Mechanism | Set on | Effect | Hard or soft |
|---|---|---|---|
nodeName | Pod | Skips the scheduler entirely; the kubelet on that node runs it | Hard, ignores taints of effect NoSchedule |
nodeSelector | Pod | Node must have all listed labels | Hard |
| Node affinity | Pod | Label expressions with In, NotIn, Exists, DoesNotExist, Gt, Lt | required... hard, preferred... weighted |
| Pod affinity / anti-affinity | Pod | Co-locate with, or keep away from, Pods matching a selector, per topologyKey | Both forms |
| Taint | Node | Repels Pods that don't tolerate it | NoSchedule, NoExecute hard; PreferNoSchedule soft |
| Toleration | Pod | Allows (doesn't require) scheduling onto a tainted node | n/a |
| Topology spread | Pod | Caps the imbalance of matching Pods across zones or nodes | DoNotSchedule or ScheduleAnyway |
| PriorityClass | Pod | Order in the queue, and permission to preempt lower-priority Pods | n/a |
Labels, nodeSelector and node affinity
k label node worker-3 storage=nvme
k label node worker-3 storage- # remove the label
k get nodes -L storage,topology.kubernetes.io/zoneWell-known labels you can rely on: kubernetes.io/hostname, kubernetes.io/os, kubernetes.io/arch,
topology.kubernetes.io/zone, topology.kubernetes.io/region, node.kubernetes.io/instance-type.
spec:
nodeSelector:
kubernetes.io/os: linux
affinity:
nodeAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
nodeSelectorTerms: # terms are ORed
- matchExpressions: # expressions inside one term are ANDed
- key: storage
operator: In
values: ["nvme", "ssd"]
preferredDuringSchedulingIgnoredDuringExecution:
- weight: 50 # 1 to 100
preference:
matchExpressions:
- key: topology.kubernetes.io/zone
operator: In
values: ["zone-b"]- IgnoredDuringExecution: removing the label later doesn't evict running Pods.
nodeSelectorand node affinity together must both match.NotInandDoesNotExistexpress "keep off these nodes" without touching the nodes.
Exam signal
"Must run only on nodes labelled X" is nodeSelector or required node affinity. "Should prefer" is
preferredDuringScheduling.... "Only these Pods may use the GPU nodes" needs a taint on the nodes plus a
toleration and an affinity on the Pods.
Pod affinity and anti-affinity
affinity:
podAntiAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
- labelSelector:
matchLabels:
app: sentry-cache
topologyKey: kubernetes.io/hostname # at most one sentry-cache Pod per node
podAffinity:
preferredDuringSchedulingIgnoredDuringExecution:
- weight: 80
podAffinityTerm:
labelSelector:
matchLabels:
app: sentry-api
topologyKey: topology.kubernetes.io/zone # same zone as the API if possibletopologyKeydefines the "place": a node (hostname) or a zone. It must be a label on the nodes.- Required anti-affinity per node with more replicas than nodes leaves the extras
Pending. Use the preferred form, or topology spread, when replicas can outnumber nodes.
Taints and tolerations
k taint nodes worker-4 dedicated=analytics:NoSchedule
k taint nodes worker-4 dedicated=analytics:NoSchedule- # remove (note the trailing -)
k describe node worker-4 | grep -i taintstolerations:
- key: dedicated
operator: Equal # Exists matches any value; omit value with Exists
value: analytics
effect: NoSchedule
- key: node.kubernetes.io/unreachable
operator: Exists
effect: NoExecute
tolerationSeconds: 60 # leave an unreachable node after 60 s instead of 300 s| Effect | New Pods without toleration | Running Pods without toleration |
|---|---|---|
PreferNoSchedule | Avoided if possible | Stay |
NoSchedule | Not scheduled | Stay |
NoExecute | Not scheduled | Evicted (after tolerationSeconds, if set in their toleration) |
- kubeadm taints control plane nodes with
node-role.kubernetes.io/control-plane:NoSchedule. DaemonSets that must run there need that toleration. - The node controller adds
node.kubernetes.io/not-readyandnode.kubernetes.io/unreachableNoExecute taints to failing nodes. Every Pod gets default tolerations for them of 300 seconds, which is why Pods take about 5 minutes to move off a dead node.
A toleration as a node selector
A toleration only permits a Pod onto a tainted node; the scheduler is still free to put it anywhere else. For "these Pods must run on the analytics nodes and nothing else may", you need both halves: the taint keeps others off, and node affinity (or nodeSelector) pins your Pods on.
Topology spread constraints
topologySpreadConstraints:
- maxSkew: 1
topologyKey: topology.kubernetes.io/zone
whenUnsatisfiable: DoNotSchedule # or ScheduleAnyway (soft)
labelSelector:
matchLabels:
app: sentry-api
matchLabelKeys: ["pod-template-hash"] # count only Pods of the same rollout revisionmaxSkewis the largest allowed difference in matching Pods between the fullest and emptiest domain.- Unlike required anti-affinity, it scales past one Pod per domain: 9 replicas in 3 zones go 3/3/3.
minDomains(withDoNotSchedule) treats missing domains as having zero Pods, forcing spread across at least that many.
PriorityClass and preemption
k create priorityclass payments-critical --value=100000 --description="customer-facing payments"
k create priorityclass batch-low --value=100 --preemption-policy=Neverspec:
priorityClassName: payments-critical- Higher
valueschedules first. When no node fits, the scheduler may evict lower-priority Pods to make room (preemption). PDBs are respected on a best-effort basis only. preemptionPolicy: Neverkeeps the queue priority but never evicts others.globalDefault: trueon one class sets the priority for Pods without a class; otherwise they get 0.- Built-in
system-cluster-criticalandsystem-node-critical(about 2 billion) are reserved for cluster components.
Scenarios
The taint keeps all other Pods off the GPU nodes, the toleration lets trainer onto them, and the nodeSelector keeps trainer from landing anywhere else. A taint with toleration alone still lets trainer run on other nodes. A nodeSelector alone lets other Pods onto the GPU nodes. Anti-affinity against Pods doesn't reserve nodes.
Required anti-affinity per hostname allows at most one Pod per node, so 5 replicas can't fit on 3 nodes. A spread constraint with maxSkew 1 places them 2/2/1. Using zones makes it stricter if there are fewer zones than nodes. The unschedulable toleration is for cordoned nodes, and preemption can't satisfy an anti-affinity rule against the Deployment's own Pods.
NoSchedule blocks new placements but leaves running Pods alone. NoExecute evicts Pods that don't tolerate it,
and kubectl drain cordons and evicts while honouring PDBs, which is the gentler choice. Operators belong to
tolerations, not taints. The scheduler doesn't re-place running Pods, and there is no default toleration for
arbitrary NoSchedule taints.
Drill
kubectl config use-context drill-w6. In namespace harbor:
- Label node
worker-2withtier=edgeand taint ittier=edge:NoSchedule. - Create Deployment
beacon(imagenginx:1.27-alpine, 2 replicas) whose Pods run only ontier=edgenodes. - Create PriorityClass
beacon-highwith value50000and use it forbeacon.
kubectl config use-context drill-w6
k label node worker-2 tier=edge
k taint node worker-2 tier=edge:NoSchedule
k create priorityclass beacon-high --value=50000
k create deploy beacon -n harbor --image=nginx:1.27-alpine --replicas=2 --dry-run=client -o yaml > beacon.yamlAdd to spec.template.spec in beacon.yaml:
priorityClassName: beacon-high
tolerations:
- key: tier
operator: Equal
value: edge
effect: NoSchedule
affinity:
nodeAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
nodeSelectorTerms:
- matchExpressions:
- key: tier
operator: In
values: ["edge"]k apply -f beacon.yaml
k get pods -n harbor -l app=beacon -o wide # NODE is worker-2 for both
k get pods -n harbor -l app=beacon -o jsonpath='{.items[*].spec.priority}{"\n"}' # 50000 50000A nodeSelector: {tier: edge} would pin the Pods just as well as the affinity block.
Further reading
Resources, QoS and quotas
Container requests and limits, CPU throttling and OOM kills, QoS classes and eviction order, LimitRange defaults and bounds, ResourceQuota on compute and object counts, and how admission rejects Pods.
Domain 3 · Services and networking
20% of the exam. How Pods reach each other, how traffic gets in through Services, Gateway API and Ingress, how NetworkPolicies fence it, and how CoreDNS turns names into IPs.