Asterrr's Handbook

Domain 5 · Troubleshooting

30% of the exam, the heaviest domain. Fixing NotReady nodes, broken control plane components, noisy and failing workloads, missing logs and Services that don't answer.

Domain 5 tasks hand you something that is already broken and ask you to make it work again. A node is NotReady, kubectl can't reach the API server, a Deployment never gets Pods, a container restarts every minute, or a Service returns nothing. The points come from finding the one broken thing fast and proving the fix, not from rebuilding the cluster.

CompetencyWhat it's really askingPages
5.1 Troubleshoot clusters and nodesWhy a node is NotReady or why Pods don't start on it: kubelet, container runtime, certificates, kubeconfig, node conditionsTroubleshooting nodes, Troubleshooting applications
5.2 Troubleshoot cluster componentsA static Pod manifest with a typo, an API server that won't start, a scheduler or controller manager that's downTroubleshooting the control plane
5.3 Monitor cluster and application resource usagekubectl top, metrics-server, events, and finding the Pod or node that eats the most CPU or memoryResource usage
5.4 Manage and evaluate container output streamskubectl logs with the right flags, logs of a crashed container, logs on the node when the API is goneContainer logs, Troubleshooting applications
5.5 Troubleshoot services and networkingServices without endpoints, wrong ports, DNS that doesn't resolve, kube-proxy, NetworkPolicies that block too muchTroubleshooting networking

Triage: symptom to page

Start from what you can see, not from what you suspect. The first command that fails tells you which layer to open.

The commands you'll run on almost every task

k get nodes -o wide                       # node status, kubelet version, runtime
k get pods -A -o wide | grep -v Running   # everything that isn't healthy, with its node
k describe pod <pod> -n <ns>              # Events at the bottom: the reason, in plain words
k events -n <ns> --types=Warning          # warnings, oldest first
k logs <pod> -n <ns> --previous           # output of the container that just crashed
ssh <node>; sudo journalctl -u kubelet -e # node side: why the kubelet is unhappy

What connects Domain 5 to the rest of the exam:

Fixing the symptom instead of the cause

Deleting a crashing Pod, restarting the kubelet or scaling a Deployment to zero and back often makes the status look green for a few seconds. The grader checks the state minutes later. Read the event or log line that names the cause, fix that (the image tag, the manifest path, the selector, the certificate), then watch the object settle with k get -w before you move on.

On this page