Domain 5 · Troubleshooting
30% of the exam, the heaviest domain. Fixing NotReady nodes, broken control plane components, noisy and failing workloads, missing logs and Services that don't answer.
Domain 5 tasks hand you something that is already broken and ask you to make it work again. A node is
NotReady, kubectl can't reach the API server, a Deployment never gets Pods, a container restarts every
minute, or a Service returns nothing. The points come from finding the one broken thing fast and proving the
fix, not from rebuilding the cluster.
| Competency | What it's really asking | Pages |
|---|---|---|
| 5.1 Troubleshoot clusters and nodes | Why a node is NotReady or why Pods don't start on it: kubelet, container runtime, certificates, kubeconfig, node conditions | Troubleshooting nodes, Troubleshooting applications |
| 5.2 Troubleshoot cluster components | A static Pod manifest with a typo, an API server that won't start, a scheduler or controller manager that's down | Troubleshooting the control plane |
| 5.3 Monitor cluster and application resource usage | kubectl top, metrics-server, events, and finding the Pod or node that eats the most CPU or memory | Resource usage |
| 5.4 Manage and evaluate container output streams | kubectl logs with the right flags, logs of a crashed container, logs on the node when the API is gone | Container logs, Troubleshooting applications |
| 5.5 Troubleshoot services and networking | Services without endpoints, wrong ports, DNS that doesn't resolve, kube-proxy, NetworkPolicies that block too much | Troubleshooting networking |
Triage: symptom to page
Start from what you can see, not from what you suspect. The first command that fails tells you which layer to open.
The commands you'll run on almost every task
k get nodes -o wide # node status, kubelet version, runtime
k get pods -A -o wide | grep -v Running # everything that isn't healthy, with its node
k describe pod <pod> -n <ns> # Events at the bottom: the reason, in plain words
k events -n <ns> --types=Warning # warnings, oldest first
k logs <pod> -n <ns> --previous # output of the container that just crashed
ssh <node>; sudo journalctl -u kubelet -e # node side: why the kubelet is unhappyWhat connects Domain 5 to the rest of the exam:
- Most broken things here were built in other domains. Static Pods, kubeadm paths and certificates come from kubeadm install and etcd backup and restore.
- Pending Pods are often a scheduling problem: see Pod placement and Resources and quotas.
- Network fixes lean on Services, CoreDNS and Network policies.
Fixing the symptom instead of the cause
Deleting a crashing Pod, restarting the kubelet or scaling a Deployment to zero and back often makes the status
look green for a few seconds. The grader checks the state minutes later. Read the event or log line that names
the cause, fix that (the image tag, the manifest path, the selector, the certificate), then watch the object
settle with k get -w before you move on.
StorageClasses and dynamic provisioning
How a StorageClass and its CSI driver create PVs on demand, setting the default class, Immediate vs WaitForFirstConsumer binding, allowVolumeExpansion and growing a claim, and which fields you can't change later.
Troubleshooting nodes
Why a node goes NotReady and how to bring it back, from node conditions and leases to the kubelet, the container runtime, kubeconfig and certificates, using systemctl and journalctl.