Asterrr's Handbook

Scheduling

How kube-scheduler filters and scores nodes, and how resource requests, nodeSelector, node and Pod affinity, taints and tolerations, and priority steer where Pods land.

Exam tasks: 1.3 (Scheduling: the scheduler, resource requests, placement constraints)

The decision: which mechanism makes a Pod land on (or stay off) a particular set of nodes, and is it set on the Pod or on the node?

How the scheduler decides

  • Filtering removes nodes that can't host the Pod: not enough unrequested CPU or memory, a taint the Pod doesn't tolerate, a nodeSelector or required affinity that doesn't match, a host port already in use.
  • Scoring ranks the survivors: spreading replicas across nodes and zones, preferred affinity, how evenly resources would be used, whether the image is already on the node. The highest score wins; ties are broken at random.
  • Each step is a plugin in the scheduling framework (PreFilter, Filter, PostFilter, Score, Reserve, Permit, Bind and others), so the default scheduler can be tuned with profiles, and you can run extra schedulers and pick one with spec.schedulerName.
  • If no node fits, the Pod stays Pending with a FailedScheduling event explaining why. The scheduler tries again when the cluster changes, for example when a node joins or another Pod is deleted.

Exam signal

"Pending with 0/5 nodes are available" is a scheduling problem: requests too large, an untolerated taint, or a selector no node matches. It's never the kubelet's or the image's fault, because no node has been chosen yet.

Requests and limits

RequestLimit
Used byThe scheduler, to find a node with roomThe kubelet and runtime, at run time
MeaningCapacity reserved for the containerThe most it may use
CPU over the valuen/aThrottled, keeps running
Memory over the valuen/aContainer is OOMKilled
  • The scheduler adds up requests, not actual usage. A node can be nearly idle and still refuse a Pod because its capacity is already reserved.
  • CPU is measured in cores (500m is half a core); memory in bytes (256Mi).
  • The combination gives a Pod its QoS class: Guaranteed (requests equal limits for every container), Burstable (some requests or limits set), BestEffort (none). Under memory pressure the kubelet evicts BestEffort first and Guaranteed last.
  • Quotas and defaults per namespace (ResourceQuota, LimitRange) are covered in CKA resources and quotas.

Limits drive scheduling

Only requests count when the scheduler checks whether a Pod fits. A Pod with a tiny request and a huge limit schedules easily and may then compete for memory at run time.

Steering placement

MechanismSet onEffectStrength
nodeNamePodSkip the scheduler, bind directlyAbsolute, and bypasses most checks
nodeSelectorPodOnly nodes with all these labelsHard
Node affinityPodExpressions over node labels (In, NotIn, Exists, Gt...)Hard (required...) or soft (preferred...)
Pod affinity / anti-affinityPodNear or away from other Pods with given labels, per topology key (node, zone)Hard or soft
Topology spread constraintsPodEven spread across zones or nodes, with a maxSkewHard or soft
TaintNodeRepels Pods that don't tolerate itDepends on effect
TolerationPodAllows (doesn't force) a Pod onto a tainted nodePermission only
  • The IgnoredDuringExecution suffix on affinity rules means a running Pod isn't moved if node labels change later.
  • Attract with selectors and affinity, which live on the Pod. Repel with taints, which live on the node.

Taints and tolerations

# kubectl taint nodes gpu-node-1 accelerator=nvidia:NoSchedule
apiVersion: v1
kind: Pod
metadata:
  name: render-worker
spec:
  tolerations:
    - key: accelerator
      operator: Equal
      value: nvidia
      effect: NoSchedule
  nodeSelector:
    accelerator: nvidia       # the node also carries this label
  containers:
    - name: render
      image: registry.example.com/render-worker:1.8
EffectNew Pods without a tolerationPods already running
NoScheduleNot scheduledStay
PreferNoScheduleAvoided if possibleStay
NoExecuteNot scheduledEvicted (after tolerationSeconds, if set)
  • kubeadm taints control plane nodes with node-role.kubernetes.io/control-plane:NoSchedule, which is why your Pods don't land there.
  • The node controller adds taints such as node.kubernetes.io/not-ready and node.kubernetes.io/unreachable (NoExecute) to failing nodes, which is how Pods get evicted from a dead node.

A toleration sends the Pod to the tainted node

A toleration only permits placement. To make dedicated nodes run only a certain workload and make that workload run only there, combine a taint on the nodes with a toleration plus a nodeSelector or node affinity on the Pod, as in the example above.

Priority and preemption

  • A PriorityClass gives Pods an integer priority. When a high-priority Pod can't be scheduled, the scheduler may preempt (evict) lower-priority Pods to make room.
  • Built-in classes system-cluster-critical and system-node-critical protect components like CoreDNS.
  • Hands-on detail for all of these: CKA pod placement.

Scenarios

Scenario
Velora's cluster has four nodes reserved for a machine learning team. Other teams' Pods must not be placed on them. What should be configured on the four nodes?
Scenario
A Pod requests 6 GiB of memory and stays Pending. Every node has 8 GiB, but other Pods already request 4 GiB on each node, even though they only use about 1 GiB. Why isn't the Pod scheduled?
Scenario
Which statement about the scheduler is correct?

Further reading

On this page