Troubleshooting applications
Reading Pod status and events to fix Pending, ImagePullBackOff, CreateContainerConfigError, CrashLoopBackOff, OOMKilled and not-ready Pods, plus kubectl exec and kubectl debug for containers without a shell.
Exam tasks: 5.1 (troubleshoot workloads on clusters and nodes), 5.4 (evaluate container output)
The decision: what does the Pod's status say failed (scheduling, image, configuration, the process itself, or a probe), and which single field fixes it?
Read the status, then the events
k get pod -n atlas -o wide
k describe pod lumen-6f7c9 -n atlas | tail -20 # Events: the reason in words
k get pod lumen-6f7c9 -n atlas -o jsonpath='{.status.containerStatuses[0].lastState}'
k logs lumen-6f7c9 -n atlas --previousStatus by status
| Status | What it means | Typical fix |
|---|---|---|
Pending, event FailedScheduling: 0/3 nodes are available: 3 Insufficient cpu | Requests don't fit any node | Lower requests, free capacity, or add a node |
Pending, untolerated taint / didn't match Pod's node affinity/selector | Placement rules exclude every node | Fix the nodeSelector, label the node, add a toleration (Pod placement) |
Pending, pod has unbound immediate PersistentVolumeClaims | PVC can't bind | Fix the StorageClass name, size or access mode (Persistent volumes) |
ContainerCreating, FailedMount ... configmap "x" not found | A volume source doesn't exist | Create the ConfigMap or Secret, or fix its name |
ErrImagePull then ImagePullBackOff | Wrong image or tag, private registry without credentials | Fix the image with k set image, add imagePullSecrets |
CreateContainerConfigError | env / envFrom points at a missing ConfigMap, Secret or key | Create it, or fix the key name |
RunContainerError / StartError, exec: "x": executable file not found | command names a binary that isn't in the image | Fix command / args |
CrashLoopBackOff | The process keeps exiting; kubelet waits longer between restarts | Read logs --previous and the exit code |
OOMKilled (in lastState.terminated.reason), exit code 137 | Container went over its memory limit | Raise the limit, or fix the app |
Running but READY 0/1 | Readiness probe fails; Pod gets no Service traffic | Fix probe port, path or initialDelaySeconds, or the app |
Terminating forever | Finalizer or unreachable node | Check metadata.finalizers; k delete pod x --force --grace-period=0 as a last resort |
ImagePullBackOff and the wrong fix
Deleting the Pod only gets you a new Pod with the same bad image. Fix the template (Deployment, StatefulSet)
with k set image deploy/lumen lumen=registry.example.com/lumen:2.4.1 -n atlas or k edit, then watch the
rollout. Also check for a typo in the registry host, not just the tag.
Exit codes worth knowing
| Code | Meaning |
|---|---|
| 0 | Process finished. With restartPolicy: Always, a Deployment still restarts it, and a short-lived command looks like a crash loop |
| 1 | Application error. Read the logs |
| 126 / 127 | Command not executable / not found |
| 137 | Killed with SIGKILL: OOM kill, or a liveness probe failure that outlived the grace period |
| 139 | Segmentation fault |
| 143 | SIGTERM: normal shutdown, for example during a rollout |
OOMKilled, it hit its memory limit.containerStatuses holding the previous instance's exit code, reason and times.kubectl debug node/... mounts the node's root filesystem in the debug Pod.Exam signal
A liveness probe that is too aggressive also produces restarts and CrashLoopBackOff. If the logs look healthy
and describe shows Liveness probe failed events followed by Killing, fix the probe (port, path,
initialDelaySeconds, or add a startupProbe), not the app.
Getting inside: exec and debug
k exec -it lumen-6f7c9 -n atlas -- sh # image has a shell
k exec lumen-6f7c9 -n atlas -c lumen -- cat /etc/lumen/config.yamlMinimal and distroless images have no shell. Use kubectl debug:
# ephemeral container in the running Pod, sharing the app container's process namespace
k debug -it lumen-6f7c9 -n atlas --image=busybox:1.37 --target=lumen
# copy of the Pod, with the app container's command replaced by a shell, so it doesn't crash on start
k debug lumen-6f7c9 -n atlas -it --copy-to=lumen-dbg --container=lumen -- sh
# copy with a different image for every container
k debug lumen-6f7c9 -n atlas --copy-to=lumen-dbg --set-image='*=busybox:1.37'
# a privileged-ish Pod on a node, with the node's filesystem at /host
k debug node/worker-2 -it --image=busybox:1.37
chroot /host # inside it: commands now see the node's filesystem as /- Ephemeral containers can't be removed once added, and they don't restart. They're fine for a look around.
--profilepicks the security settings of the debug container; the default isgeneral. Usenetadminfor network tools that needNET_ADMIN,sysadminfor full privileges.- Clean up copies and node debug Pods afterwards:
k delete pod lumen-dbg -n atlas.
Scenarios
Keys are case-sensitive, so DB_HOST and db_host are different keys. Either make the ConfigMap provide the key the Pod asks for or point the reference at the existing key. Recreating Pods hits the same error, and the image and memory have nothing to do with configuration lookup.
OOMKilled with exit code 137 means the container's cgroup hit its memory limit, regardless of free memory on the node. A node eviction shows the Pod as Evicted with a node pressure message instead. Liveness kills show Liveness probe failed and Killing events, not reason OOMKilled. PID pressure is a node condition.
An ephemeral container brings its own tools (busybox), joins the running Pod, and with --target shares the app
container's process namespace, so ps shows the app and /proc/1/root exposes its filesystem. The copy option
starts a new Pod and tries to run sh from the distroless image, which has none. bash is just as absent as sh.
A node debug Pod sees the node, not inside the container.
Drill
kubectl config use-context lab-apps. Deployment beacon in namespace harbor should run 3 ready replicas,
but its Pods keep restarting. Find the cause and fix it in the Deployment. Don't change the image.
k get pods -n harbor -l app=beacon # CrashLoopBackOff, restarts climbing
k describe pod -n harbor -l app=beacon | grep -A6 'Last State'
# Last State: Terminated Reason: OOMKilled Exit Code: 137
k get deploy beacon -n harbor -o jsonpath='{.spec.template.spec.containers[0].resources}'
# {"limits":{"memory":"32Mi"},"requests":{"memory":"16Mi"}}
k set resources deploy beacon -n harbor -c beacon --requests=memory=128Mi --limits=memory=256Mi
k rollout status deploy beacon -n harborIf the reason had been Error with exit code 1, the fix would come from k logs --previous instead (a missing
env var, a bad argument). If describe showed Liveness probe failed, the probe would be the thing to edit.
Verify:
k get deploy beacon -n harbor # 3/3 READY
k get pods -n harbor -l app=beacon # RESTARTS stays at 0 for a minuteFurther reading
Container logs
Reading container stdout and stderr with kubectl logs (previous instances, multi-container Pods, selectors, time windows), where the logs live on the node, rotation, and streaming file logs through a sidecar.
Troubleshooting networking
Tracing a failed request hop by hop, from DNS through the Service and its EndpointSlices to kube-proxy, NetworkPolicies and the CNI, with the test Pod commands that prove each hop.