Troubleshooting the control plane
Finding and fixing a broken kube-apiserver, kube-scheduler, kube-controller-manager or etcd on a kubeadm cluster, through static Pod manifests, crictl, Pod log files and health endpoints.
Exam tasks: 5.2 (troubleshoot cluster components)
The decision: which control plane component is broken, and can you still use kubectl to find out, or do
you have to work from the node with crictl and log files?
Which component is down?
Each component fails in a recognisable way. Match the symptom before you open a manifest.
| Symptom | Broken component | Why |
|---|---|---|
kubectl says "connection to the server …:6443 was refused" or times out | kube-apiserver (or etcd behind it) | Nothing answers on 6443, so every client fails |
New Pods stay Pending, describe shows no events, NODE is <none> | kube-scheduler | Nobody sets spec.nodeName, so there is not even a FailedScheduling event |
k create deploy succeeds but no ReplicaSet or Pods appear; scaling does nothing | kube-controller-manager | Controllers (Deployment, ReplicaSet, node lifecycle, endpoints) aren't reconciling |
API server restarts every few seconds, logs mention 2379 or "context deadline exceeded" | etcd | API server can't reach its datastore |
| Pods are scheduled to a node but never start | kubelet on that node | See Troubleshooting nodes |
Static Pods: how kubeadm runs the control plane
- The kubelet on each control plane node watches
staticPodPath(kubeadm sets/etc/kubernetes/manifests) and runs every manifest there:kube-apiserver.yaml,kube-controller-manager.yaml,kube-scheduler.yaml,etcd.yaml. - The API server shows a read-only mirror Pod for each, named
<component>-<node-name>, for examplekube-scheduler-cp-1. Deleting the mirror Pod withkubectldoesn't stop it; the kubelet recreates it. - To change a component you edit the file on disk. The kubelet notices, kills the old container and starts a new one. Give it 20 to 60 seconds.
- To stop a component, move its manifest out of the directory. Moving it back starts it again.
Backups inside the manifests directory
cp kube-apiserver.yaml kube-apiserver.yaml.bak inside /etc/kubernetes/manifests leaves a second manifest
for a Pod with the same name, and the kubelet may run the old copy instead of your edit. Keep backups somewhere
else, such as /root/ or /tmp/.
When the API server is down
kubectl is useless, so go to the node. Logs live in three places; use whichever has content.
ssh cp-1
sudo crictl ps -a | grep kube-apiserver # Exited? how many attempts?
sudo crictl logs <container-id> 2>&1 | tail -20 # last words of the failed container
sudo ls /var/log/pods/ | grep apiserver # files survive container removal
sudo tail -30 /var/log/pods/kube-system_kube-apiserver-cp-1_*/kube-apiserver/*.log
sudo journalctl -u kubelet --no-pager | grep -i apiserver | tail # YAML parse errors land here- No container at all usually means the kubelet can't parse the manifest (bad indentation, a tab, a missing
quote). The error is in
journalctl -u kubelet, not in a container log. - Container exits immediately usually means a bad flag or path: an unknown flag, a cert file that doesn't
exist, a wrong
--etcd-serversURL. The container log names it. - Paths in flags must exist inside the container. Check the manifest's
volumeMountsandhostPathvolumes: a cert moved to a new directory also needs a new mount.
Exam signal
Diff against a healthy reference instead of reading 40 flags. On a multi-node control plane, diff the
manifest with another control plane node's copy. On a single node, compare with the backup you (or the task)
made, or with the flags in the kube-apiserver reference page. Usually the broken one is the only value that
looks off: 2380 where 2379 belongs, /etc/kubernetes/pki/ca.crtx, --authorization-mode=Nod,RBAC.
Scheduler, controller manager and etcd
k get pods -n kube-system -o wide
k describe pod kube-scheduler-cp-1 -n kube-system # Events: probe failures, image errors
k logs kube-scheduler-cp-1 -n kube-system # flag errors, kubeconfig problems- Scheduler and controller manager each use their own kubeconfig (
/etc/kubernetes/scheduler.conf,/etc/kubernetes/controller-manager.conf). A wrong path in--kubeconfigor a broken file stops them from talking to the API server. - Liveness probes in the manifests hit
/healthzon ports 10259 (scheduler) and 10257 (controller manager) over HTTPS. A changed--secure-portwithout matching probe port gives restart loops. - etcd uses
/etc/kubernetes/pki/etcd/certs and keeps data in thehostPathbehind--data-dir. A wrong data dir after a restore is a classic break; see etcd backup and restore.
Health checks
k get --raw='/readyz?verbose' # every API server readiness check, one per line
k get --raw='/livez' # ok
curl -k https://localhost:10259/healthz # scheduler, from the control plane node
curl -k https://localhost:10257/healthz # controller managerLegacy: use the API server's /livez and /readyz endpoints and the component Pods' status instead
kubectl get componentstatuses (cs) has been deprecated since v1.19. It may still print something, but it
doesn't reflect HA control planes or etcd accurately. Don't rely on it to prove a fix.
Scenarios
No ReplicaSet means the Deployment controller isn't running, and that lives in the controller manager. It's a static Pod, so the fix is in its manifest on disk. Recreating the Deployment changes nothing while the controller is down. Deleting the mirror Pod only makes the kubelet recreate the same broken container. The scheduler places Pods; it can't create a ReplicaSet.
No container means the kubelet never got as far as starting one, typically because the YAML is invalid. The kubelet logs the parse error. There's no container log to read, and kubectl can't reach a dead API server.
Taints, insufficient resources and affinity all produce a FailedScheduling event from the scheduler. No event at all means no scheduler looked at the Pod. A quota rejection happens at creation time, so the Pod wouldn't exist. Kubelets only act after a node is assigned.
Drill
kubectl config use-context lab-cp. kubectl on this cluster stopped working after a change to the API
server. Restore it without running kubeadm init again. The control plane node is cp-orion.
k get nodes # connection refused
ssh cp-orion
sudo crictl ps -a | grep kube-apiserver
sudo crictl logs $(sudo crictl ps -a --name kube-apiserver -q | head -1) 2>&1 | tail -5
# example finding: "open /etc/kubernetes/pki/apiserver.crtt: no such file or directory"
sudo cp /etc/kubernetes/manifests/kube-apiserver.yaml /root/kube-apiserver.yaml.before
sudo vi /etc/kubernetes/manifests/kube-apiserver.yaml # fix the flag the log namesIf crictl ps -a shows nothing, read sudo journalctl -u kubelet --no-pager | tail -30 for a YAML error
instead, and fix the indentation the message points to.
Wait for the kubelet to restart it, then verify:
watch sudo crictl ps --name kube-apiserver # Running, attempt count stops rising
exit
k get --raw='/readyz' # ok
k get pods -n kube-system # kube-apiserver-cp-orion RunningFurther reading
Troubleshooting nodes
Why a node goes NotReady and how to bring it back, from node conditions and leases to the kubelet, the container runtime, kubeconfig and certificates, using systemctl and journalctl.
Resource usage
Measuring what nodes, Pods and containers actually consume with metrics-server and kubectl top, telling usage apart from requests, and using events to explain evictions and pressure.