Asterrr's Handbook

Troubleshooting the control plane

Finding and fixing a broken kube-apiserver, kube-scheduler, kube-controller-manager or etcd on a kubeadm cluster, through static Pod manifests, crictl, Pod log files and health endpoints.

Exam tasks: 5.2 (troubleshoot cluster components)

The decision: which control plane component is broken, and can you still use kubectl to find out, or do you have to work from the node with crictl and log files?

Which component is down?

Each component fails in a recognisable way. Match the symptom before you open a manifest.

SymptomBroken componentWhy
kubectl says "connection to the server …:6443 was refused" or times outkube-apiserver (or etcd behind it)Nothing answers on 6443, so every client fails
New Pods stay Pending, describe shows no events, NODE is <none>kube-schedulerNobody sets spec.nodeName, so there is not even a FailedScheduling event
k create deploy succeeds but no ReplicaSet or Pods appear; scaling does nothingkube-controller-managerControllers (Deployment, ReplicaSet, node lifecycle, endpoints) aren't reconciling
API server restarts every few seconds, logs mention 2379 or "context deadline exceeded"etcdAPI server can't reach its datastore
Pods are scheduled to a node but never startkubelet on that nodeSee Troubleshooting nodes

Static Pods: how kubeadm runs the control plane

  • The kubelet on each control plane node watches staticPodPath (kubeadm sets /etc/kubernetes/manifests) and runs every manifest there: kube-apiserver.yaml, kube-controller-manager.yaml, kube-scheduler.yaml, etcd.yaml.
  • The API server shows a read-only mirror Pod for each, named <component>-<node-name>, for example kube-scheduler-cp-1. Deleting the mirror Pod with kubectl doesn't stop it; the kubelet recreates it.
  • To change a component you edit the file on disk. The kubelet notices, kills the old container and starts a new one. Give it 20 to 60 seconds.
  • To stop a component, move its manifest out of the directory. Moving it back starts it again.

Backups inside the manifests directory

cp kube-apiserver.yaml kube-apiserver.yaml.bak inside /etc/kubernetes/manifests leaves a second manifest for a Pod with the same name, and the kubelet may run the old copy instead of your edit. Keep backups somewhere else, such as /root/ or /tmp/.

When the API server is down

kubectl is useless, so go to the node. Logs live in three places; use whichever has content.

ssh cp-1
sudo crictl ps -a | grep kube-apiserver          # Exited? how many attempts?
sudo crictl logs <container-id> 2>&1 | tail -20  # last words of the failed container
sudo ls /var/log/pods/ | grep apiserver          # files survive container removal
sudo tail -30 /var/log/pods/kube-system_kube-apiserver-cp-1_*/kube-apiserver/*.log
sudo journalctl -u kubelet --no-pager | grep -i apiserver | tail   # YAML parse errors land here
  • No container at all usually means the kubelet can't parse the manifest (bad indentation, a tab, a missing quote). The error is in journalctl -u kubelet, not in a container log.
  • Container exits immediately usually means a bad flag or path: an unknown flag, a cert file that doesn't exist, a wrong --etcd-servers URL. The container log names it.
  • Paths in flags must exist inside the container. Check the manifest's volumeMounts and hostPath volumes: a cert moved to a new directory also needs a new mount.

Exam signal

Diff against a healthy reference instead of reading 40 flags. On a multi-node control plane, diff the manifest with another control plane node's copy. On a single node, compare with the backup you (or the task) made, or with the flags in the kube-apiserver reference page. Usually the broken one is the only value that looks off: 2380 where 2379 belongs, /etc/kubernetes/pki/ca.crtx, --authorization-mode=Nod,RBAC.

Scheduler, controller manager and etcd

k get pods -n kube-system -o wide
k describe pod kube-scheduler-cp-1 -n kube-system   # Events: probe failures, image errors
k logs kube-scheduler-cp-1 -n kube-system           # flag errors, kubeconfig problems
  • Scheduler and controller manager each use their own kubeconfig (/etc/kubernetes/scheduler.conf, /etc/kubernetes/controller-manager.conf). A wrong path in --kubeconfig or a broken file stops them from talking to the API server.
  • Liveness probes in the manifests hit /healthz on ports 10259 (scheduler) and 10257 (controller manager) over HTTPS. A changed --secure-port without matching probe port gives restart loops.
  • etcd uses /etc/kubernetes/pki/etcd/ certs and keeps data in the hostPath behind --data-dir. A wrong data dir after a restore is a classic break; see etcd backup and restore.

Health checks

k get --raw='/readyz?verbose'       # every API server readiness check, one per line
k get --raw='/livez'                # ok
curl -k https://localhost:10259/healthz   # scheduler, from the control plane node
curl -k https://localhost:10257/healthz   # controller manager

Legacy: use the API server's /livez and /readyz endpoints and the component Pods' status instead

kubectl get componentstatuses (cs) has been deprecated since v1.19. It may still print something, but it doesn't reflect HA control planes or etcd accurately. Don't rely on it to prove a fix.

/etc/kubernetes/manifests
kubeadm's static Pod directory for control plane components.
6443
kube-apiserver secure port on kubeadm clusters.
2379 / 2380
etcd client port (API server connects here) and peer port (etcd members).
10259 / 10257
Secure health ports of kube-scheduler and kube-controller-manager.
/var/log/pods
Per-Pod log directories, readable even when the API server is down.

Scenarios

Scenario
Developers report that a new Deployment named invoice-gen shows 0/2 ready. `k get rs -l app=invoice-gen` returns nothing, and `k get pods -n kube-system` shows kube-controller-manager-cp-1 in CrashLoopBackOff. Which action addresses the root cause?
Scenario
After someone edited /etc/kubernetes/manifests/kube-apiserver.yaml, kubectl returns 'connection refused'. On the node, `crictl ps -a | grep apiserver` shows no container at all, not even an exited one. Where is the error most likely reported?
Scenario
A new Pod named probe-runner stays Pending. `k describe pod probe-runner` shows Node: none and an empty Events list. Nodes are Ready with plenty of free CPU and memory. What should you investigate?

Drill

kubectl config use-context lab-cp. kubectl on this cluster stopped working after a change to the API server. Restore it without running kubeadm init again. The control plane node is cp-orion.

Further reading

On this page