Debugging applications
Debugging a running application on Kubernetes with logs, events, kubectl exec, port-forward, kubectl debug and ephemeral containers, and reading why a Deployment rollout is stuck.
Exam tasks: 3.2 (debugging: inspecting running workloads, ephemeral containers, failing rollouts)
The decision: which kubectl command gets you the evidence fastest: what the app printed, what the cluster
recorded, what's inside the container, or what happens when you call it directly?
Pick the tool
| Command | What it shows or does |
|---|---|
kubectl describe pod web-7d9f | Spec, container states, restart counts, probe failures and the Events at the bottom |
kubectl events -n shop or kubectl get events --sort-by=.lastTimestamp | Cluster-recorded events: scheduling, image pulls, probe failures, OOM kills |
kubectl logs web-7d9f -c api | stdout and stderr of one container |
kubectl logs web-7d9f --previous | Logs from the last crashed instance, the ones you need for CrashLoopBackOff |
kubectl logs deploy/web -f | Follow logs from a Pod of a Deployment |
kubectl exec -it web-7d9f -- sh | A shell inside a running container |
kubectl port-forward svc/web 8080:80 | Tunnel local port 8080 to the Service's port 80 through the API server |
kubectl debug -it web-7d9f --image=busybox:1.37 --target=api | Add an ephemeral container to a running Pod |
kubectl debug node/worker-2 -it --image=busybox:1.37 | A Pod on that node with the host filesystem mounted at /host |
kubectl top pod | Live CPU and memory, if metrics-server is installed |
Exam signal
Events are short-lived: by default the API server keeps them for about one hour. For a failure from last
night, events are gone, so the answer is your logging or monitoring system, not kubectl get events.
Logs and events first
- Kubernetes doesn't store application logs long-term.
kubectl logsreads what the kubelet still has on the node for that container. When the Pod is deleted, those logs go with it. - Apps should log to stdout and stderr. A node-level agent (Fluent Bit, for example) ships them to a central
store. Logs written to a file inside the container don't show in
kubectl logs. - Events are API objects that controllers and the kubelet write:
FailedScheduling,Pulling,BackOff,Unhealthy,Killing. They explain what the cluster did, not what the app did.
exec, port-forward and debug
kubectl execneeds the target container to be running and to contain the binary you call. Minimal and distroless images often have no shell, soexec ... -- shfails.kubectl port-forwardis for testing a Pod or Service from your machine without creating a NodePort, LoadBalancer or Ingress. It's a debugging tunnel, not a way to expose an app to users.kubectl debugworks in three modes:- Ephemeral container added to the running Pod (
--targetshares the process namespace of one container, so you can see its processes). - Copy of the Pod (
--copy-to=web-debug), optionally with a changed image or command, so you can poke at a crashing app without touching the original. - Node debugging (
kubectl debug node/<name>), which runs a Pod on that node with the host's root filesystem mounted.
- Ephemeral container added to the running Pod (
Ephemeral containers
- They're added through the Pod's
ephemeralcontainerssubresource, which is whykubectl debugcan change a Pod that's otherwise immutable. - They bring your tools (shell,
curl,nslookup,tcpdump) into the Pod while the app image stays small. - RBAC still applies: you need permission on
pods/ephemeralcontainers(and onpods/execforexec).
Rebuild the image with debugging tools
Adding a shell and network tools to the production image so you can exec in widens the attack surface of every
running copy. On the exam, the cloud native answer for "the image has no shell" is kubectl debug with an
ephemeral container.
Reading a failing rollout
- A rolling update never removes old Pods faster than new ones become available, so a bad version usually leaves you with a mix: old Pods still serving, new Pods stuck.
- If no progress happens for
progressDeadlineSeconds(600 by default), the Deployment'sProgressingcondition turnsFalsewith reasonProgressDeadlineExceeded. Kubernetes reports the failure but doesn't roll back on its own. - Usual causes in the new Pods:
ImagePullBackOff(wrong tag or missing pull secret),CrashLoopBackOff(app fails at startup, checklogs --previous), readiness probe failing (wrong path or port), orPending(not enough CPU or memory to schedule the surge Pods). kubectl rollout undo deployment/webreturns to the previous revision. Under GitOps, revert the commit instead, or the agent will reapply the broken version.
| Symptom in the new Pods | First command | Typical cause |
|---|---|---|
ImagePullBackOff | describe pod (Events) | Tag doesn't exist, private registry without imagePullSecrets |
CrashLoopBackOff | logs --previous | Missing config or env var, app exits at startup |
Running but 0/1 Ready | describe pod (probe failures) | Readiness probe path or port wrong, dependency down |
Pending | describe pod (FailedScheduling) | Requests too large, taints, no matching node |
OOMKilled in last state | describe pod | Memory limit too low |
Scenarios
An ephemeral container brings its own tools into the existing Pod without a restart, and --target lets it see
the app container's processes. exec fails because there's no shell. You can't add regular containers to a
running Pod; changing the spec through a Deployment would replace the Pod. Port-forwarding tests the app from
outside, not DNS or processes from inside.
CrashLoopBackOff means the container keeps starting and exiting, and --previous shows the output of the
instance that just crashed, usually the error itself. Events confirm the back-off but not the app's reason. top
shows resource use, and rollout history lists revisions without saying why one fails.
kubectl port-forward svc/<name> <local>:<remote> tunnels traffic through the API server and creates nothing in
the cluster. Exposing a NodePort or adding an Ingress changes cluster objects and widens access. Node debugging
starts a Pod on a node and isn't meant for calling a Service from your laptop.
Further reading
Packaging and rollouts
Helm charts, values and releases vs Kustomize bases and overlays, Deployment rolling updates and rollback, blue/green and canary releases, and progressive delivery with Argo Rollouts and Flagger.
Domain 4 · Cloud native architecture
12% of the exam. Observability signals and tools, the principles behind cloud native design and autoscaling, the CNCF projects and their maturity, and how the community governs Kubernetes.