Troubleshooting Pods
Pod phases, container states and their reasons, events, kubectl describe and logs, and how to map Pending, ImagePullBackOff, CrashLoopBackOff and OOMKilled to their usual causes for KCNA.
Exam tasks: 2.3 (troubleshooting)
The decision: a Pod isn't doing its job. Given its status, which part of the system failed (scheduling, image pull, the app itself, resources, networking), and which command shows you the evidence?
Phases vs states vs reasons
A Pod reports status at two levels. Questions often mix them up on purpose.
| Level | Values | Set by |
|---|---|---|
Pod phase (status.phase) | Pending, Running, Succeeded, Failed, Unknown | The kubelet and control plane, a summary of the whole Pod |
| Container state | Waiting, Running, Terminated | The kubelet, per container |
| Reason | ContainerCreating, ImagePullBackOff, CrashLoopBackOff, OOMKilled, Completed, Error... | Explains a Waiting or Terminated state |
- The
STATUScolumn ofkubectl get podsshows the most useful of these, often a reason, not the phase.CrashLoopBackOffis never a phase; the Pod's phase isRunningwhile one of its containers waits to restart. - Pending: accepted by the API server but not all containers have started (not scheduled yet, or images still pulling).
- Succeeded / Failed: all containers have terminated, with success, or at least one with failure and no more restarts. Typical for Jobs.
- Unknown: the control plane lost contact with the node.
The troubleshooting order
| Command | What it shows |
|---|---|
kubectl get pods -o wide | Status, restarts, Pod IP, which node |
kubectl describe pod <name> | Conditions, each container's state and last state with exit code, and recent events |
kubectl logs <pod> | The container's stdout and stderr |
kubectl logs <pod> --previous | Logs from the previous run of a container that crashed and restarted |
kubectl logs <pod> -c <container> | One container in a multi-container Pod |
kubectl get events -n <ns> --sort-by=.lastTimestamp | All recent events in a namespace, oldest first |
kubectl top pod | Live CPU and memory use (needs metrics-server) |
- Events are API objects recorded by the scheduler, kubelet and controllers. They are kept for a limited time (one hour by default), so look soon after the failure.
kubectl logsonly works once the container has started. For a Pod stuck before that, events are your only source.
Exam signal
"The container keeps restarting and the current logs are empty" means read kubectl logs --previous. "The Pod
never started, where do you look?" means kubectl describe pod and its Events section.
Common failures
| Status / reason | Usual cause | Where the evidence is |
|---|---|---|
| Pending (not scheduled) | No node has enough requested CPU or memory, taints without tolerations, node selector or affinity matches nothing, unbound PVC | Event FailedScheduling from the scheduler |
| ContainerCreating (stuck) | Volume can't attach or mount, CNI can't set up the network, Secret or ConfigMap volume missing | Events from the kubelet (FailedMount) |
| ErrImagePull then ImagePullBackOff | Wrong image name or tag, private registry without an imagePullSecret, registry unreachable | Events (Failed to pull image) |
| CreateContainerConfigError | Env var references a missing ConfigMap or Secret key | Events |
| CrashLoopBackOff | The process starts and exits again and again: app error, bad command, missing config, failing liveness probe | kubectl logs --previous, exit code in describe |
| OOMKilled (exit code 137) | The container used more memory than its limit | Last state in describe |
| Running but 0/1 Ready | Readiness probe failing, so the Pod is removed from Service endpoints | Events (Readiness probe failed) |
| Completed | The process exited 0. Normal for Jobs, wrong for a long-running Deployment | Check the container command |
- BackOff means the kubelet is waiting before the next try, with the delay doubling up to five minutes.
- Exit code 137 is 128 + 9 (SIGKILL). 143 is 128 + 15 (SIGTERM, graceful stop). 1 is a generic app error.
Raise the CPU limit for OOMKilled
OOMKilled is about memory, never CPU. A container over its CPU limit is throttled (slowed), not killed. The fix is a higher memory limit or a fix to the memory leak.
Restart it
"Delete the Pod so it gets recreated" rarely fixes a CrashLoopBackOff or ImagePullBackOff. The replacement has
the same spec and fails the same way. Fix the spec (image, config, limits), and the controller rolls out working
Pods.
Beyond the Pod
- Service has no traffic: check that its selector matches Pod labels and that Pods are Ready
(
kubectl get endpointslices -l kubernetes.io/service-name=<svc>). An empty list means no matching Ready Pods. - Name doesn't resolve: check CoreDNS Pods in
kube-systemand the Service name and namespace. - Node NotReady: the kubelet stopped reporting (crashed, out of disk or memory, network or CNI down). Pods on
it eventually show
Unknownand are replaced elsewhere if a controller owns them. - For interactive tools (
exec,port-forward,kubectl debugwith ephemeral containers), see Debugging. For hands-on depth, see the CKA page Troubleshooting applications.
Scenarios
The scheduler places Pods by their requests. No node has enough unreserved memory, so the Pod can't be scheduled and stays Pending. Image pull problems happen after scheduling and show ImagePullBackOff. OOMKilled needs a running container. If kubelets were down, nodes would be NotReady and the message would differ.
The current container may have just restarted and logged nothing yet; --previous shows the run that crashed.
CrashLoopBackOff is a container reason, not a phase. The scheduler already placed the Pod. The image clearly
exists because the container has started seven times.
OOMKilled means the container went over its memory limit and the kernel killed it. Only memory changes help. CPU overuse causes throttling, not kills. A readiness probe controls traffic, not memory. Deployments only allow restartPolicy Always, and stopping restarts wouldn't fix the memory use anyway.
Further reading
Kubernetes security
The 4Cs of cloud native security, the API request path (authentication, authorization, admission), RBAC and ServiceAccounts, Pod Security Standards, Secrets, and image and supply chain security for KCNA.
Kubernetes storage
Ephemeral and persistent volumes, PersistentVolumes and PersistentVolumeClaims, access modes and reclaim policies, StorageClasses and dynamic provisioning, CSI drivers, and StatefulSet storage for KCNA.