Asterrr's Handbook

Debugging applications

Debugging a running application on Kubernetes with logs, events, kubectl exec, port-forward, kubectl debug and ephemeral containers, and reading why a Deployment rollout is stuck.

Exam tasks: 3.2 (debugging: inspecting running workloads, ephemeral containers, failing rollouts)

The decision: which kubectl command gets you the evidence fastest: what the app printed, what the cluster recorded, what's inside the container, or what happens when you call it directly?

Pick the tool

CommandWhat it shows or does
kubectl describe pod web-7d9fSpec, container states, restart counts, probe failures and the Events at the bottom
kubectl events -n shop or kubectl get events --sort-by=.lastTimestampCluster-recorded events: scheduling, image pulls, probe failures, OOM kills
kubectl logs web-7d9f -c apistdout and stderr of one container
kubectl logs web-7d9f --previousLogs from the last crashed instance, the ones you need for CrashLoopBackOff
kubectl logs deploy/web -fFollow logs from a Pod of a Deployment
kubectl exec -it web-7d9f -- shA shell inside a running container
kubectl port-forward svc/web 8080:80Tunnel local port 8080 to the Service's port 80 through the API server
kubectl debug -it web-7d9f --image=busybox:1.37 --target=apiAdd an ephemeral container to a running Pod
kubectl debug node/worker-2 -it --image=busybox:1.37A Pod on that node with the host filesystem mounted at /host
kubectl top podLive CPU and memory, if metrics-server is installed

Exam signal

Events are short-lived: by default the API server keeps them for about one hour. For a failure from last night, events are gone, so the answer is your logging or monitoring system, not kubectl get events.

Logs and events first

  • Kubernetes doesn't store application logs long-term. kubectl logs reads what the kubelet still has on the node for that container. When the Pod is deleted, those logs go with it.
  • Apps should log to stdout and stderr. A node-level agent (Fluent Bit, for example) ships them to a central store. Logs written to a file inside the container don't show in kubectl logs.
  • Events are API objects that controllers and the kubelet write: FailedScheduling, Pulling, BackOff, Unhealthy, Killing. They explain what the cluster did, not what the app did.

exec, port-forward and debug

  • kubectl exec needs the target container to be running and to contain the binary you call. Minimal and distroless images often have no shell, so exec ... -- sh fails.
  • kubectl port-forward is for testing a Pod or Service from your machine without creating a NodePort, LoadBalancer or Ingress. It's a debugging tunnel, not a way to expose an app to users.
  • kubectl debug works in three modes:
    • Ephemeral container added to the running Pod (--target shares the process namespace of one container, so you can see its processes).
    • Copy of the Pod (--copy-to=web-debug), optionally with a changed image or command, so you can poke at a crashing app without touching the original.
    • Node debugging (kubectl debug node/<name>), which runs a Pod on that node with the host's root filesystem mounted.

Ephemeral containers

Stable
Ephemeral containers have been GA since Kubernetes 1.25.
No restart
They're never restarted, and you can't remove one from a Pod once it's added.
No ports, no probes
They can't declare ports, probes or resources. They exist for inspection, not for serving.
  • They're added through the Pod's ephemeralcontainers subresource, which is why kubectl debug can change a Pod that's otherwise immutable.
  • They bring your tools (shell, curl, nslookup, tcpdump) into the Pod while the app image stays small.
  • RBAC still applies: you need permission on pods/ephemeralcontainers (and on pods/exec for exec).

Rebuild the image with debugging tools

Adding a shell and network tools to the production image so you can exec in widens the attack surface of every running copy. On the exam, the cloud native answer for "the image has no shell" is kubectl debug with an ephemeral container.

Reading a failing rollout

  • A rolling update never removes old Pods faster than new ones become available, so a bad version usually leaves you with a mix: old Pods still serving, new Pods stuck.
  • If no progress happens for progressDeadlineSeconds (600 by default), the Deployment's Progressing condition turns False with reason ProgressDeadlineExceeded. Kubernetes reports the failure but doesn't roll back on its own.
  • Usual causes in the new Pods: ImagePullBackOff (wrong tag or missing pull secret), CrashLoopBackOff (app fails at startup, check logs --previous), readiness probe failing (wrong path or port), or Pending (not enough CPU or memory to schedule the surge Pods).
  • kubectl rollout undo deployment/web returns to the previous revision. Under GitOps, revert the commit instead, or the agent will reapply the broken version.
Symptom in the new PodsFirst commandTypical cause
ImagePullBackOffdescribe pod (Events)Tag doesn't exist, private registry without imagePullSecrets
CrashLoopBackOfflogs --previousMissing config or env var, app exits at startup
Running but 0/1 Readydescribe pod (probe failures)Readiness probe path or port wrong, dependency down
Pendingdescribe pod (FailedScheduling)Requests too large, taints, no matching node
OOMKilled in last statedescribe podMemory limit too low

Scenarios

Scenario
A Pod runs a distroless image with no shell. You need to inspect its processes and test DNS from inside the Pod without restarting it. What should you do?
Scenario
After an image update, kubectl rollout status for a Deployment times out. The old Pods still serve traffic, and the new Pods show CrashLoopBackOff. Which command most directly shows why the new containers fail?
Scenario
A developer wants to send test requests from their workstation to a ClusterIP Service in a development cluster, without changing any objects in the cluster. Which command fits?

Further reading

On this page