Asterrr's Handbook

Troubleshooting applications

Reading Pod status and events to fix Pending, ImagePullBackOff, CreateContainerConfigError, CrashLoopBackOff, OOMKilled and not-ready Pods, plus kubectl exec and kubectl debug for containers without a shell.

Exam tasks: 5.1 (troubleshoot workloads on clusters and nodes), 5.4 (evaluate container output)

The decision: what does the Pod's status say failed (scheduling, image, configuration, the process itself, or a probe), and which single field fixes it?

Read the status, then the events

k get pod -n atlas -o wide
k describe pod lumen-6f7c9 -n atlas | tail -20            # Events: the reason in words
k get pod lumen-6f7c9 -n atlas -o jsonpath='{.status.containerStatuses[0].lastState}'
k logs lumen-6f7c9 -n atlas --previous

Status by status

StatusWhat it meansTypical fix
Pending, event FailedScheduling: 0/3 nodes are available: 3 Insufficient cpuRequests don't fit any nodeLower requests, free capacity, or add a node
Pending, untolerated taint / didn't match Pod's node affinity/selectorPlacement rules exclude every nodeFix the nodeSelector, label the node, add a toleration (Pod placement)
Pending, pod has unbound immediate PersistentVolumeClaimsPVC can't bindFix the StorageClass name, size or access mode (Persistent volumes)
ContainerCreating, FailedMount ... configmap "x" not foundA volume source doesn't existCreate the ConfigMap or Secret, or fix its name
ErrImagePull then ImagePullBackOffWrong image or tag, private registry without credentialsFix the image with k set image, add imagePullSecrets
CreateContainerConfigErrorenv / envFrom points at a missing ConfigMap, Secret or keyCreate it, or fix the key name
RunContainerError / StartError, exec: "x": executable file not foundcommand names a binary that isn't in the imageFix command / args
CrashLoopBackOffThe process keeps exiting; kubelet waits longer between restartsRead logs --previous and the exit code
OOMKilled (in lastState.terminated.reason), exit code 137Container went over its memory limitRaise the limit, or fix the app
Running but READY 0/1Readiness probe fails; Pod gets no Service trafficFix probe port, path or initialDelaySeconds, or the app
Terminating foreverFinalizer or unreachable nodeCheck metadata.finalizers; k delete pod x --force --grace-period=0 as a last resort

ImagePullBackOff and the wrong fix

Deleting the Pod only gets you a new Pod with the same bad image. Fix the template (Deployment, StatefulSet) with k set image deploy/lumen lumen=registry.example.com/lumen:2.4.1 -n atlas or k edit, then watch the rollout. Also check for a typo in the registry host, not just the tag.

Exit codes worth knowing

CodeMeaning
0Process finished. With restartPolicy: Always, a Deployment still restarts it, and a short-lived command looks like a crash loop
1Application error. Read the logs
126 / 127Command not executable / not found
137Killed with SIGKILL: OOM kill, or a liveness probe failure that outlived the grace period
139Segmentation fault
143SIGTERM: normal shutdown, for example during a rollout
10s → 5 min
CrashLoopBackOff delay: starts around 10 seconds, doubles each restart, capped at 5 minutes. Resets after 10 minutes of running fine.
137
Exit code of a SIGKILLed container. With reason OOMKilled, it hit its memory limit.
lastState
Field under containerStatuses holding the previous instance's exit code, reason and times.
/host
Where kubectl debug node/... mounts the node's root filesystem in the debug Pod.

Exam signal

A liveness probe that is too aggressive also produces restarts and CrashLoopBackOff. If the logs look healthy and describe shows Liveness probe failed events followed by Killing, fix the probe (port, path, initialDelaySeconds, or add a startupProbe), not the app.

Getting inside: exec and debug

k exec -it lumen-6f7c9 -n atlas -- sh                 # image has a shell
k exec lumen-6f7c9 -n atlas -c lumen -- cat /etc/lumen/config.yaml

Minimal and distroless images have no shell. Use kubectl debug:

# ephemeral container in the running Pod, sharing the app container's process namespace
k debug -it lumen-6f7c9 -n atlas --image=busybox:1.37 --target=lumen

# copy of the Pod, with the app container's command replaced by a shell, so it doesn't crash on start
k debug lumen-6f7c9 -n atlas -it --copy-to=lumen-dbg --container=lumen -- sh

# copy with a different image for every container
k debug lumen-6f7c9 -n atlas --copy-to=lumen-dbg --set-image='*=busybox:1.37'

# a privileged-ish Pod on a node, with the node's filesystem at /host
k debug node/worker-2 -it --image=busybox:1.37
chroot /host        # inside it: commands now see the node's filesystem as /
  • Ephemeral containers can't be removed once added, and they don't restart. They're fine for a look around.
  • --profile picks the security settings of the debug container; the default is general. Use netadmin for network tools that need NET_ADMIN, sysadmin for full privileges.
  • Clean up copies and node debug Pods afterwards: k delete pod lumen-dbg -n atlas.

Scenarios

Scenario
Deployment ferry in namespace docks shows 0/2 ready. Its Pods are in CreateContainerConfigError. `k describe pod` shows: Error: couldn't find key DB_HOST in ConfigMap docks/ferry-env. `k get cm ferry-env -n docks -o yaml` shows a key named db_host. What is the cleanest fix?
Scenario
Pod pulse-0 restarts every few minutes. `k describe pod pulse-0` shows Last State: Terminated, Reason: OOMKilled, Exit Code: 137, and the container has limits.memory: 64Mi. Node worker-1 has 12 GiB free. What is happening?
Scenario
Pod quartz-api runs a distroless image and is Running, but it answers every request with HTTP 500. `k exec -it quartz-api -- sh` fails with 'executable file not found'. You want to see the app's processes and read files from its filesystem without restarting it. Which command fits?

Drill

kubectl config use-context lab-apps. Deployment beacon in namespace harbor should run 3 ready replicas, but its Pods keep restarting. Find the cause and fix it in the Deployment. Don't change the image.

Further reading

On this page