Asterrr's Handbook

Troubleshooting Pods

Pod phases, container states and their reasons, events, kubectl describe and logs, and how to map Pending, ImagePullBackOff, CrashLoopBackOff and OOMKilled to their usual causes for KCNA.

Exam tasks: 2.3 (troubleshooting)

The decision: a Pod isn't doing its job. Given its status, which part of the system failed (scheduling, image pull, the app itself, resources, networking), and which command shows you the evidence?

Phases vs states vs reasons

A Pod reports status at two levels. Questions often mix them up on purpose.

LevelValuesSet by
Pod phase (status.phase)Pending, Running, Succeeded, Failed, UnknownThe kubelet and control plane, a summary of the whole Pod
Container stateWaiting, Running, TerminatedThe kubelet, per container
ReasonContainerCreating, ImagePullBackOff, CrashLoopBackOff, OOMKilled, Completed, Error...Explains a Waiting or Terminated state
  • The STATUS column of kubectl get pods shows the most useful of these, often a reason, not the phase. CrashLoopBackOff is never a phase; the Pod's phase is Running while one of its containers waits to restart.
  • Pending: accepted by the API server but not all containers have started (not scheduled yet, or images still pulling).
  • Succeeded / Failed: all containers have terminated, with success, or at least one with failure and no more restarts. Typical for Jobs.
  • Unknown: the control plane lost contact with the node.

The troubleshooting order

CommandWhat it shows
kubectl get pods -o wideStatus, restarts, Pod IP, which node
kubectl describe pod <name>Conditions, each container's state and last state with exit code, and recent events
kubectl logs <pod>The container's stdout and stderr
kubectl logs <pod> --previousLogs from the previous run of a container that crashed and restarted
kubectl logs <pod> -c <container>One container in a multi-container Pod
kubectl get events -n <ns> --sort-by=.lastTimestampAll recent events in a namespace, oldest first
kubectl top podLive CPU and memory use (needs metrics-server)
  • Events are API objects recorded by the scheduler, kubelet and controllers. They are kept for a limited time (one hour by default), so look soon after the failure.
  • kubectl logs only works once the container has started. For a Pod stuck before that, events are your only source.

Exam signal

"The container keeps restarting and the current logs are empty" means read kubectl logs --previous. "The Pod never started, where do you look?" means kubectl describe pod and its Events section.

Common failures

Status / reasonUsual causeWhere the evidence is
Pending (not scheduled)No node has enough requested CPU or memory, taints without tolerations, node selector or affinity matches nothing, unbound PVCEvent FailedScheduling from the scheduler
ContainerCreating (stuck)Volume can't attach or mount, CNI can't set up the network, Secret or ConfigMap volume missingEvents from the kubelet (FailedMount)
ErrImagePull then ImagePullBackOffWrong image name or tag, private registry without an imagePullSecret, registry unreachableEvents (Failed to pull image)
CreateContainerConfigErrorEnv var references a missing ConfigMap or Secret keyEvents
CrashLoopBackOffThe process starts and exits again and again: app error, bad command, missing config, failing liveness probekubectl logs --previous, exit code in describe
OOMKilled (exit code 137)The container used more memory than its limitLast state in describe
Running but 0/1 ReadyReadiness probe failing, so the Pod is removed from Service endpointsEvents (Readiness probe failed)
CompletedThe process exited 0. Normal for Jobs, wrong for a long-running DeploymentCheck the container command
  • BackOff means the kubelet is waiting before the next try, with the delay doubling up to five minutes.
  • Exit code 137 is 128 + 9 (SIGKILL). 143 is 128 + 15 (SIGTERM, graceful stop). 1 is a generic app error.

Raise the CPU limit for OOMKilled

OOMKilled is about memory, never CPU. A container over its CPU limit is throttled (slowed), not killed. The fix is a higher memory limit or a fix to the memory leak.

Restart it

"Delete the Pod so it gets recreated" rarely fixes a CrashLoopBackOff or ImagePullBackOff. The replacement has the same spec and fails the same way. Fix the spec (image, config, limits), and the controller rolls out working Pods.

Beyond the Pod

  • Service has no traffic: check that its selector matches Pod labels and that Pods are Ready (kubectl get endpointslices -l kubernetes.io/service-name=<svc>). An empty list means no matching Ready Pods.
  • Name doesn't resolve: check CoreDNS Pods in kube-system and the Service name and namespace.
  • Node NotReady: the kubelet stopped reporting (crashed, out of disk or memory, network or CNI down). Pods on it eventually show Unknown and are replaced elsewhere if a controller owns them.
  • For interactive tools (exec, port-forward, kubectl debug with ephemeral containers), see Debugging. For hands-on depth, see the CKA page Troubleshooting applications.
5 phases
Pending, Running, Succeeded, Failed, Unknown.
3 states
Waiting, Running, Terminated, per container.
137
Exit code for SIGKILL, often OOMKilled.
1 hour
Default event retention.
5 minutes
Maximum restart back-off delay.

Scenarios

Scenario
A new Pod named report-builder has stayed in Pending for ten minutes. kubectl describe shows the event FailedScheduling: 0/4 nodes are available: 4 Insufficient memory. What is the cause?
Scenario
A Pod shows STATUS CrashLoopBackOff with 7 restarts. Running kubectl logs on it returns nothing useful. What should you check next?
Scenario
A container in a Deployment is repeatedly terminated with reason OOMKilled and exit code 137. Which change addresses the cause?

Further reading

On this page