Scheduling
How kube-scheduler filters and scores nodes, and how resource requests, nodeSelector, node and Pod affinity, taints and tolerations, and priority steer where Pods land.
Exam tasks: 1.3 (Scheduling: the scheduler, resource requests, placement constraints)
The decision: which mechanism makes a Pod land on (or stay off) a particular set of nodes, and is it set on the Pod or on the node?
How the scheduler decides
- Filtering removes nodes that can't host the Pod: not enough unrequested CPU or memory, a taint the Pod
doesn't tolerate, a
nodeSelectoror required affinity that doesn't match, a host port already in use. - Scoring ranks the survivors: spreading replicas across nodes and zones, preferred affinity, how evenly resources would be used, whether the image is already on the node. The highest score wins; ties are broken at random.
- Each step is a plugin in the scheduling framework (PreFilter, Filter, PostFilter, Score, Reserve, Permit,
Bind and others), so the default scheduler can be tuned with profiles, and you can run extra schedulers and
pick one with
spec.schedulerName. - If no node fits, the Pod stays Pending with a
FailedSchedulingevent explaining why. The scheduler tries again when the cluster changes, for example when a node joins or another Pod is deleted.
Exam signal
"Pending with 0/5 nodes are available" is a scheduling problem: requests too large, an untolerated taint, or a selector no node matches. It's never the kubelet's or the image's fault, because no node has been chosen yet.
Requests and limits
| Request | Limit | |
|---|---|---|
| Used by | The scheduler, to find a node with room | The kubelet and runtime, at run time |
| Meaning | Capacity reserved for the container | The most it may use |
| CPU over the value | n/a | Throttled, keeps running |
| Memory over the value | n/a | Container is OOMKilled |
- The scheduler adds up requests, not actual usage. A node can be nearly idle and still refuse a Pod because its capacity is already reserved.
- CPU is measured in cores (
500mis half a core); memory in bytes (256Mi). - The combination gives a Pod its QoS class: Guaranteed (requests equal limits for every container), Burstable (some requests or limits set), BestEffort (none). Under memory pressure the kubelet evicts BestEffort first and Guaranteed last.
- Quotas and defaults per namespace (ResourceQuota, LimitRange) are covered in CKA resources and quotas.
Limits drive scheduling
Only requests count when the scheduler checks whether a Pod fits. A Pod with a tiny request and a huge limit schedules easily and may then compete for memory at run time.
Steering placement
| Mechanism | Set on | Effect | Strength |
|---|---|---|---|
nodeName | Pod | Skip the scheduler, bind directly | Absolute, and bypasses most checks |
nodeSelector | Pod | Only nodes with all these labels | Hard |
| Node affinity | Pod | Expressions over node labels (In, NotIn, Exists, Gt...) | Hard (required...) or soft (preferred...) |
| Pod affinity / anti-affinity | Pod | Near or away from other Pods with given labels, per topology key (node, zone) | Hard or soft |
| Topology spread constraints | Pod | Even spread across zones or nodes, with a maxSkew | Hard or soft |
| Taint | Node | Repels Pods that don't tolerate it | Depends on effect |
| Toleration | Pod | Allows (doesn't force) a Pod onto a tainted node | Permission only |
- The
IgnoredDuringExecutionsuffix on affinity rules means a running Pod isn't moved if node labels change later. - Attract with selectors and affinity, which live on the Pod. Repel with taints, which live on the node.
Taints and tolerations
# kubectl taint nodes gpu-node-1 accelerator=nvidia:NoSchedule
apiVersion: v1
kind: Pod
metadata:
name: render-worker
spec:
tolerations:
- key: accelerator
operator: Equal
value: nvidia
effect: NoSchedule
nodeSelector:
accelerator: nvidia # the node also carries this label
containers:
- name: render
image: registry.example.com/render-worker:1.8| Effect | New Pods without a toleration | Pods already running |
|---|---|---|
NoSchedule | Not scheduled | Stay |
PreferNoSchedule | Avoided if possible | Stay |
NoExecute | Not scheduled | Evicted (after tolerationSeconds, if set) |
- kubeadm taints control plane nodes with
node-role.kubernetes.io/control-plane:NoSchedule, which is why your Pods don't land there. - The node controller adds taints such as
node.kubernetes.io/not-readyandnode.kubernetes.io/unreachable(NoExecute) to failing nodes, which is how Pods get evicted from a dead node.
A toleration sends the Pod to the tainted node
A toleration only permits placement. To make dedicated nodes run only a certain workload and make that workload run only there, combine a taint on the nodes with a toleration plus a nodeSelector or node affinity on the Pod, as in the example above.
Priority and preemption
- A PriorityClass gives Pods an integer priority. When a high-priority Pod can't be scheduled, the scheduler may preempt (evict) lower-priority Pods to make room.
- Built-in classes
system-cluster-criticalandsystem-node-criticalprotect components like CoreDNS. - Hands-on detail for all of these: CKA pod placement.
Scenarios
Taints are set on nodes and repel every Pod that lacks a matching toleration, so other teams are kept off without changing their manifests. A label only helps Pods that opt in. Tolerations belong on Pods, not nodes. PriorityClass affects preemption, not which nodes are allowed.
Scheduling is based on the sum of requests, not real usage: each node has only 4 GiB unrequested. The kubelet never sees the Pod because no node was chosen. A missing limit doesn't block scheduling, and nothing in the question says the nodes are under memory pressure.
Filter, score, bind is the scheduler's whole job. The kubelet starts containers. Rules ending in IgnoredDuringExecution don't move running Pods, and the default scheduler uses requests, not live metrics.
Further reading
The API and kubectl
API groups and versions, object structure, declarative vs imperative management, everyday kubectl, RBAC basics, and which tool builds which kind of cluster.
Containers and runtimes
What a container really is, images and layers, tags and digests, the three OCI specs, registries, the CRI, containerd and CRI-O, low-level runtimes, and sandboxed runtimes with RuntimeClass.