Troubleshooting nodes
Why a node goes NotReady and how to bring it back, from node conditions and leases to the kubelet, the container runtime, kubeconfig and certificates, using systemctl and journalctl.
Exam tasks: 5.1 (troubleshoot clusters and nodes)
The decision: the node is NotReady (or Pods on it never start). Is it the kubelet, the container runtime,
the kubelet's credentials, or the node running out of something?
How a node becomes NotReady
The kubelet on each node renews a Lease in kube-node-lease and posts node status. The node lifecycle
controller in kube-controller-manager watches both.
| What you see | Most likely cause | Where to look |
|---|---|---|
Ready is Unknown, "Kubelet stopped posting node status" | kubelet stopped, crashed, can't reach the API server, or the node is down | systemctl status kubelet, journalctl -u kubelet |
Ready is False, "container runtime is down" or "PLEG is not healthy" | containerd or CRI-O stopped or hung | systemctl status containerd, crictl info |
Ready is False, "network plugin not ready: cni config uninitialized" | No CNI installed, or its config is missing on this node | /etc/cni/net.d/, CNI DaemonSet Pod on that node |
MemoryPressure, DiskPressure, PIDPressure True | Node is short on memory, disk or process IDs; kubelet evicts Pods | k describe node, df -h, free -m |
Ready but SchedulingDisabled | Someone ran kubectl cordon or drain | k uncordon <node> |
Node missing from k get nodes entirely | kubelet never registered: wrong kubeconfig, expired bootstrap token, wrong API server address | journalctl -u kubelet |
--node-monitor-grace-period before the controller marks a silent node Unknown.tolerationSeconds added to Pods for the not-ready and unreachable taints, so eviction starts about 5 minutes later.kubelet-client-current.pem.Step 1: read the node from the API
k get nodes -o wide
k describe node worker-2 | sed -n '/Conditions/,/Addresses/p' # conditions and their messages
k describe node worker-2 | grep -A5 TaintsThe Message column of the conditions usually names the broken piece. Copy it into your head before you SSH.
Step 2: check the kubelet on the node
ssh worker-2
sudo systemctl status kubelet # active (running)? enabled? restarting in a loop?
sudo journalctl -u kubelet -e # jump to the end; -f to follow; --since "10 min ago"
sudo systemctl cat kubelet # unit file plus the kubeadm drop-in, shows which flags and files loadCommon kubelet failures and fixes:
- Stopped or disabled.
sudo systemctl enable --now kubelet. A task that says "make it survive reboots" wantsenable, not juststart. - Bad config file. The log shows a parse error or an unknown field in
/var/lib/kubelet/config.yaml. Fix the file, thensudo systemctl restart kubelet. - Wrong binary path or flag in the systemd drop-in or
/var/lib/kubelet/kubeadm-flags.env. After editing a unit file runsudo systemctl daemon-reloadbefore the restart. - Cgroup driver mismatch. The kubelet's
cgroupDrivermust match the runtime's (SystemdCgroup = truein containerd'sconfig.tomlfor thesystemddriver). - Can't reach the API server. "connection refused" or "x509" errors to the
server:in/etc/kubernetes/kubelet.conf. Check the address and port, and that the CA in that file matches the cluster.
Exam signal
journalctl -u kubelet is long. Search it: sudo journalctl -u kubelet --no-pager | grep -iE 'error|fail' | tail -20.
The last error before the restart loop is the one that matters.
Swap and cgroup v1
Older material says "turn off swap or the kubelet won't start". Current kubelets can run with swap when
failSwapOn: false is set, but the default is still true, so on a lab node whose log complains about swap,
sudo swapoff -a (and removing the swap line from /etc/fstab) is the quick fix. Separately, from v1.35 the kubelet won't start on a cgroup v1 host by default. If the log says
cgroup v1 is unsupported, the node OS needs cgroup v2; that isn't a kubelet flag you should flip in an exam.
Step 3: check the container runtime
sudo systemctl status containerd
sudo crictl info | head -20 # runtime ready? network ready?
sudo crictl ps -a # containers on this node, including exited ones
sudo crictl pods # Pod sandboxescrictlreads its endpoint from/etc/crictl.yaml, for exampleruntime-endpoint: unix:///run/containerd/containerd.sock. A wrong endpoint looks like "the runtime is down" even when it isn't.- If containerd is stopped:
sudo systemctl enable --now containerd, then give the kubelet a few seconds.
Legacy: use containerd or CRI-O through the CRI, inspected with crictl instead
dockershim was removed in v1.24, so docker ps on a node and systemctl status docker are no longer how you
check the runtime. v1.35 is also the last release to support containerd 1.x; plan for containerd 2.
Step 4: credentials and certificates
- kubeadm sets the kubelet up for client certificate rotation: the current cert is a symlink,
/var/lib/kubelet/pki/kubelet-client-current.pem. If it expired (a node left off for months), the log shows x509 "certificate has expired" errors. - Check:
sudo openssl x509 -noout -enddate -in /var/lib/kubelet/pki/kubelet-client-current.pem. - On a control plane node,
sudo kubeadm certs check-expirationlists every kubeadm-managed cert. - Kubelet serving certs may need CSR approval if
serverTLSBootstrapis on:k get csrthenk certificate approve <name>.
Step 5: resource pressure
DiskPressure: the kubelet evicts Pods and garbage-collects images. Find the culprit withdf -handsudo du -sh /var/lib/containerd /var/log/pods.MemoryPressure: comparefree -mwithkubectl top pods -A --sort-by=memoryon that node (see Resource usage).- Eviction thresholds come from
evictionHardin the kubelet config, for examplememory.available: "100Mi".
Scenarios
Unknown conditions mean the API server stopped hearing from the kubelet, so the kubelet itself (stopped, crashing, or unable to reach the API) is the first suspect. A CNI problem shows Ready=False with a network message, not Unknown. The controller manager is what correctly marked the node; it's working. The unreachable taint is a consequence, not a cause.
systemd caches unit definitions, so an edited unit or drop-in has no effect until daemon-reload; the restart
then launches the corrected binary. Re-joining or recreating the Node object is unnecessary because the node's
identity and certificates are intact.
The runtime is up but has no CNI config. CNI plugins are normally installed per node by a DaemonSet that writes into /etc/cni/net.d and /opt/cni/bin, so check why its Pod isn't running on this node (taints, image pull, crash). Restarting containerd doesn't create a config. Cordoning and certificates produce different messages.
Drill
kubectl config use-context lab-nodes. Node edge-worker is NotReady. Bring it back to Ready so that it
stays Ready after a reboot. Don't recreate the node.
k get nodes # edge-worker NotReady
k describe node edge-worker | grep -A8 Conditions # Unknown: kubelet stopped posting
ssh edge-worker
sudo systemctl status kubelet # inactive (dead), disabled
sudo journalctl -u kubelet -e --no-pager | tail -20If the log shows nothing after the last stop, the kubelet was simply stopped and disabled:
sudo systemctl enable --now kubeletIf it starts and immediately exits, read the last error. Typical findings and fixes:
# runtime down
sudo systemctl enable --now containerd
# typo in the config file
sudo vi /var/lib/kubelet/config.yaml && sudo systemctl restart kubelet
# edited unit or drop-in
sudo systemctl daemon-reload # systemd re-reads unit files
sudo systemctl restart kubelet # only now does the new ExecStart apply
exitVerify from the base host:
k get node edge-worker -w # Ready within a few seconds
ssh edge-worker 'systemctl is-enabled kubelet containerd' # both: enabledFurther reading
Domain 5 · Troubleshooting
30% of the exam, the heaviest domain. Fixing NotReady nodes, broken control plane components, noisy and failing workloads, missing logs and Services that don't answer.
Troubleshooting the control plane
Finding and fixing a broken kube-apiserver, kube-scheduler, kube-controller-manager or etcd on a kubeadm cluster, through static Pod manifests, crictl, Pod log files and health endpoints.