Asterrr's Handbook

Self-healing workloads

Choosing between ReplicaSet, Deployment, DaemonSet, StatefulSet, Job and CronJob, restart policies and backoff, liveness, readiness and startup probes, native sidecar containers, and PodDisruptionBudgets.

Exam tasks: 2.4 (understand the primitives used to create robust, self-healing application deployments)

The decision: which controller should own these Pods, how does Kubernetes know a container is broken or not ready yet, and how do you stop voluntary disruptions such as a node drain from taking the app down?

Pick the controller

ControllerReplaces failed PodsIdentityUpdate strategiesTypical use
DeploymentYes, through its ReplicaSetRandom suffix, interchangeableRollingUpdate, RecreateStateless services
ReplicaSetYesRandom suffixNone: template changes affect only new PodsRarely by hand; Deployments own them
StatefulSetYes, with the same name and volumename-0, name-1, stable DNS through a headless ServiceRollingUpdate (with partition), OnDeleteDatabases, quorum systems
DaemonSetYes, one per matching nodePer nodeRollingUpdate, OnDeleteLog shippers, node agents, CNI
JobRetries until completions or backoffLimitOptional completion indexn/aBatch tasks, migrations
CronJobCreates Jobs on a schedulen/an/aBackups, reports

Restart policy and backoff

  • restartPolicy is Pod-wide: Always (default, required for Deployment, StatefulSet, DaemonSet Pods), OnFailure or Never (the only values a Job allows).
  • The kubelet restarts failed containers in place on the same node, with exponential backoff: 10 s, 20 s, 40 s and so on, capped at 5 minutes. That waiting state is CrashLoopBackOff. The delay resets after a container runs cleanly for 10 minutes.
  • Rescheduling to another node is not the kubelet's job. If a node dies, the owning controller creates a replacement Pod, which the scheduler places elsewhere. A bare Pod with no controller is simply gone.

Probes

ProbeFailure meansKubernetes then
startupProbeThe app hasn't finished startingKeeps waiting; liveness and readiness don't run until it succeeds. Fails for good after failureThreshold × periodSeconds, then restarts the container
livenessProbeThe app is stuck or deadRestarts the container
readinessProbeThe app can't serve right nowMarks the Pod not Ready and removes it from Service endpoints. No restart
containers:
- name: quay
  image: registry.local/quay-api:7.0
  ports:
  - containerPort: 9090
  startupProbe:
    httpGet: { path: /healthz, port: 9090 }
    periodSeconds: 5
    failureThreshold: 24          # up to 120 s to boot
  livenessProbe:
    httpGet: { path: /healthz, port: 9090 }
    periodSeconds: 10
    failureThreshold: 3
  readinessProbe:
    tcpSocket: { port: 9090 }
    periodSeconds: 5
  • Mechanisms: httpGet (2xx or 3xx passes), tcpSocket (port accepts), exec (exit code 0), grpc (gRPC health checking protocol).
  • Defaults: periodSeconds 10, timeoutSeconds 1, failureThreshold 3, successThreshold 1, initialDelaySeconds 0.
  • Prefer a startupProbe over a long initialDelaySeconds for slow starters: it waits only as long as the app actually needs.

A liveness probe that checks dependencies

If the liveness probe calls the database and the database blips, every replica fails liveness at once and is restarted together, turning a brief dependency outage into a full one. Liveness should test only the process itself. Put dependency checks in readiness, which just takes the Pod out of rotation.

Native sidecar containers

A sidecar is an init container with restartPolicy: Always. Stable since v1.33.

spec:
  initContainers:
  - name: log-forwarder
    image: fluent/fluent-bit:4.0
    restartPolicy: Always          # this line makes it a sidecar
    volumeMounts:
    - { name: logs, mountPath: /var/log/app }
  containers:
  - name: app
    image: registry.local/billing:2.2
    volumeMounts:
    - { name: logs, mountPath: /var/log/app }
  volumes:
  - name: logs
    emptyDir: {}
  • It starts before the main containers (in init order) and keeps running alongside them; later init containers start once it has started (or its startupProbe passes).
  • It's stopped after the main containers on shutdown, so it can flush logs or proxy final requests.
  • It doesn't block a Job from completing, the problem with old-style sidecars in containers.
  • It's restarted on failure regardless of the Pod's restartPolicy, and it can have probes.

Exam signal

"Add a sidecar that must be running before the app starts" or "a Job whose sidecar keeps it from finishing" both point to initContainers with restartPolicy: Always, not another entry under containers.

Jobs and CronJobs

k create job db-migrate --image=registry.local/migrator:1.4 -n ops -- /migrate --up
k create cronjob nightly-export --image=registry.local/export:3 --schedule="15 2 * * *" -n ops -- /export
k create job export-now --from=cronjob/nightly-export -n ops   # run a CronJob once, now
FieldDefaultMeaning
completions / parallelism1 / 1Successful Pods needed, and how many run at once
backoffLimit6Retries before the Job is marked Failed
activeDeadlineSecondsnoneHard time limit for the whole Job; wins over backoffLimit
ttlSecondsAfterFinishednoneDelete the finished Job (and its Pods) after this long
completionMode: IndexedNonIndexedEach Pod gets JOB_COMPLETION_INDEX
CronJob concurrencyPolicyAllowForbid skips a run if the last one is still going, Replace kills it
CronJob timeZonecontroller's zoneIANA name such as Europe/Helsinki
CronJob successfulJobsHistoryLimit / failedJobsHistoryLimit3 / 1Finished Jobs kept
  • podFailurePolicy can fail a Job immediately on a specific exit code or ignore disruptions (such as eviction during a drain) so they don't count against backoffLimit.

PodDisruptionBudgets

k create pdb quay-pdb -n api --selector=app=quay --min-available=2
# or: --max-unavailable=1
k get pdb -n api        # ALLOWED DISRUPTIONS
  • A PDB limits voluntary disruptions that go through the Eviction API: kubectl drain, cluster autoscaler scale-down, node upgrades. It doesn't stop kubectl delete pod, node crashes or OOM kills.
  • Set either minAvailable or maxUnavailable, as a number or percentage. maxUnavailable follows scaling better.
  • unhealthyPodEvictionPolicy: AlwaysAllow lets a drain evict Pods that are already not Ready, so a crash-looping app can't block node maintenance forever.

minAvailable equal to replicas

A PDB of minAvailable: 3 on a 3-replica Deployment allows zero disruptions, so kubectl drain hangs retrying the eviction. The budget must leave room for at least one Pod to go.

5 min
Cap on the CrashLoopBackOff restart delay.
10 s / 1 s / 3
Probe defaults: periodSeconds, timeoutSeconds, failureThreshold.
6
Default Job backoffLimit.
v1.33
Native sidecar containers became stable.

Scenarios

Scenario
A Java service takes 60–90 seconds to start. Its liveness probe has initialDelaySeconds 20 and the Pods restart in a loop before they ever become Ready. The team wants slow starts tolerated but hangs after startup detected within about 30 seconds. What should you change?
Scenario
`kubectl drain node-b --ignore-daemonsets` has been retrying for ten minutes with 'Cannot evict pod as it would violate the pod's disruption budget'. The Pods belong to Deployment `cache` with 2 replicas, both on node-b, and a PDB with minAvailable 2. What is the least disruptive way to finish the drain while keeping the service up?
Scenario
A Job runs a data import container and a proxy container listed together under `containers`. The import exits 0, but the Job never completes because the proxy keeps running. How do you fix it with current Kubernetes?

Drill

kubectl config use-context drill-w4. In namespace edge:

  1. Create a DaemonSet node-probe running busybox:1.36 with sleep 86400, that also runs on control plane nodes.
  2. Create a CronJob tidy that runs busybox:1.36 with echo tidy every 30 minutes, never runs two at once, and keeps only 1 successful Job. Trigger one run now as Job tidy-manual.

Further reading

On this page