Self-healing workloads
Choosing between ReplicaSet, Deployment, DaemonSet, StatefulSet, Job and CronJob, restart policies and backoff, liveness, readiness and startup probes, native sidecar containers, and PodDisruptionBudgets.
Exam tasks: 2.4 (understand the primitives used to create robust, self-healing application deployments)
The decision: which controller should own these Pods, how does Kubernetes know a container is broken or not ready yet, and how do you stop voluntary disruptions such as a node drain from taking the app down?
Pick the controller
| Controller | Replaces failed Pods | Identity | Update strategies | Typical use |
|---|---|---|---|---|
| Deployment | Yes, through its ReplicaSet | Random suffix, interchangeable | RollingUpdate, Recreate | Stateless services |
| ReplicaSet | Yes | Random suffix | None: template changes affect only new Pods | Rarely by hand; Deployments own them |
| StatefulSet | Yes, with the same name and volume | name-0, name-1, stable DNS through a headless Service | RollingUpdate (with partition), OnDelete | Databases, quorum systems |
| DaemonSet | Yes, one per matching node | Per node | RollingUpdate, OnDelete | Log shippers, node agents, CNI |
| Job | Retries until completions or backoffLimit | Optional completion index | n/a | Batch tasks, migrations |
| CronJob | Creates Jobs on a schedule | n/a | n/a | Backups, reports |
Restart policy and backoff
restartPolicyis Pod-wide:Always(default, required for Deployment, StatefulSet, DaemonSet Pods),OnFailureorNever(the only values a Job allows).- The kubelet restarts failed containers in place on the same node, with exponential backoff: 10 s, 20 s,
40 s and so on, capped at 5 minutes. That waiting state is
CrashLoopBackOff. The delay resets after a container runs cleanly for 10 minutes. - Rescheduling to another node is not the kubelet's job. If a node dies, the owning controller creates a replacement Pod, which the scheduler places elsewhere. A bare Pod with no controller is simply gone.
Probes
| Probe | Failure means | Kubernetes then |
|---|---|---|
| startupProbe | The app hasn't finished starting | Keeps waiting; liveness and readiness don't run until it succeeds. Fails for good after failureThreshold × periodSeconds, then restarts the container |
| livenessProbe | The app is stuck or dead | Restarts the container |
| readinessProbe | The app can't serve right now | Marks the Pod not Ready and removes it from Service endpoints. No restart |
containers:
- name: quay
image: registry.local/quay-api:7.0
ports:
- containerPort: 9090
startupProbe:
httpGet: { path: /healthz, port: 9090 }
periodSeconds: 5
failureThreshold: 24 # up to 120 s to boot
livenessProbe:
httpGet: { path: /healthz, port: 9090 }
periodSeconds: 10
failureThreshold: 3
readinessProbe:
tcpSocket: { port: 9090 }
periodSeconds: 5- Mechanisms:
httpGet(2xx or 3xx passes),tcpSocket(port accepts),exec(exit code 0),grpc(gRPC health checking protocol). - Defaults:
periodSeconds10,timeoutSeconds1,failureThreshold3,successThreshold1,initialDelaySeconds0. - Prefer a startupProbe over a long
initialDelaySecondsfor slow starters: it waits only as long as the app actually needs.
A liveness probe that checks dependencies
If the liveness probe calls the database and the database blips, every replica fails liveness at once and is restarted together, turning a brief dependency outage into a full one. Liveness should test only the process itself. Put dependency checks in readiness, which just takes the Pod out of rotation.
Native sidecar containers
A sidecar is an init container with restartPolicy: Always. Stable since v1.33.
spec:
initContainers:
- name: log-forwarder
image: fluent/fluent-bit:4.0
restartPolicy: Always # this line makes it a sidecar
volumeMounts:
- { name: logs, mountPath: /var/log/app }
containers:
- name: app
image: registry.local/billing:2.2
volumeMounts:
- { name: logs, mountPath: /var/log/app }
volumes:
- name: logs
emptyDir: {}- It starts before the main containers (in init order) and keeps running alongside them; later init containers start once it has started (or its startupProbe passes).
- It's stopped after the main containers on shutdown, so it can flush logs or proxy final requests.
- It doesn't block a Job from completing, the problem with old-style sidecars in
containers. - It's restarted on failure regardless of the Pod's
restartPolicy, and it can have probes.
Exam signal
"Add a sidecar that must be running before the app starts" or "a Job whose sidecar keeps it from finishing"
both point to initContainers with restartPolicy: Always, not another entry under containers.
Jobs and CronJobs
k create job db-migrate --image=registry.local/migrator:1.4 -n ops -- /migrate --up
k create cronjob nightly-export --image=registry.local/export:3 --schedule="15 2 * * *" -n ops -- /export
k create job export-now --from=cronjob/nightly-export -n ops # run a CronJob once, now| Field | Default | Meaning |
|---|---|---|
completions / parallelism | 1 / 1 | Successful Pods needed, and how many run at once |
backoffLimit | 6 | Retries before the Job is marked Failed |
activeDeadlineSeconds | none | Hard time limit for the whole Job; wins over backoffLimit |
ttlSecondsAfterFinished | none | Delete the finished Job (and its Pods) after this long |
completionMode: Indexed | NonIndexed | Each Pod gets JOB_COMPLETION_INDEX |
CronJob concurrencyPolicy | Allow | Forbid skips a run if the last one is still going, Replace kills it |
CronJob timeZone | controller's zone | IANA name such as Europe/Helsinki |
CronJob successfulJobsHistoryLimit / failedJobsHistoryLimit | 3 / 1 | Finished Jobs kept |
podFailurePolicycan fail a Job immediately on a specific exit code or ignore disruptions (such as eviction during a drain) so they don't count againstbackoffLimit.
PodDisruptionBudgets
k create pdb quay-pdb -n api --selector=app=quay --min-available=2
# or: --max-unavailable=1
k get pdb -n api # ALLOWED DISRUPTIONS- A PDB limits voluntary disruptions that go through the Eviction API:
kubectl drain, cluster autoscaler scale-down, node upgrades. It doesn't stopkubectl delete pod, node crashes or OOM kills. - Set either
minAvailableormaxUnavailable, as a number or percentage.maxUnavailablefollows scaling better. unhealthyPodEvictionPolicy: AlwaysAllowlets a drain evict Pods that are already not Ready, so a crash-looping app can't block node maintenance forever.
minAvailable equal to replicas
A PDB of minAvailable: 3 on a 3-replica Deployment allows zero disruptions, so kubectl drain hangs
retrying the eviction. The budget must leave room for at least one Pod to go.
Scenarios
The startupProbe allows up to 120 seconds to boot and holds liveness back until it passes; after that, liveness
detects a hang within about 30 seconds. A 120-second initial delay also delays hang detection after every
restart. Readiness alone never restarts a hung container, and Deployments don't allow restartPolicy: Never.
With minAvailable equal to replicas no eviction is ever allowed. Adding a replica elsewhere and allowing one unavailable lets the drain proceed while capacity stays up. Deleting the PDB or bypassing eviction drops both Pods at once, and force-deleting has the same effect plus no graceful shutdown.
A native sidecar doesn't count toward Job completion and is stopped after the main container exits. activeDeadlineSeconds marks the Job Failed. restartPolicy Never doesn't stop a running container, and a failing liveness probe would just restart the proxy.
Drill
kubectl config use-context drill-w4. In namespace edge:
- Create a DaemonSet
node-proberunningbusybox:1.36withsleep 86400, that also runs on control plane nodes. - Create a CronJob
tidythat runsbusybox:1.36withecho tidyevery 30 minutes, never runs two at once, and keeps only 1 successful Job. Trigger one run now as Jobtidy-manual.
There's no kubectl create daemonset, so start from a Deployment and edit it.
kubectl config use-context drill-w4
k create deploy node-probe --image=busybox:1.36 -n edge --dry-run=client -o yaml -- sleep 86400 > ds.yamlIn ds.yaml: change kind to DaemonSet, delete replicas, strategy and status, and add the toleration:
apiVersion: apps/v1
kind: DaemonSet
metadata:
name: node-probe
namespace: edge
spec:
selector:
matchLabels:
app: node-probe
template:
metadata:
labels:
app: node-probe
spec:
tolerations:
- key: node-role.kubernetes.io/control-plane
operator: Exists
effect: NoSchedule
containers:
- name: busybox
image: busybox:1.36
command: ["sleep", "86400"]k apply -f ds.yaml
k get ds node-probe -n edge # DESIRED equals the number of nodes
k get pods -n edge -o wide -l app=node-probe
k create cronjob tidy -n edge --image=busybox:1.36 --schedule="*/30 * * * *" \
--dry-run=client -o yaml -- echo tidy > tidy.yaml
# add under spec: concurrencyPolicy: Forbid and successfulJobsHistoryLimit: 1
k apply -f tidy.yaml
k create job tidy-manual --from=cronjob/tidy -n edge
k get cronjob tidy -n edge -o jsonpath='{.spec.concurrencyPolicy} {.spec.successfulJobsHistoryLimit}{"\n"}'
k logs job/tidy-manual -n edge # tidyFurther reading
Workload autoscaling
Manual scaling, HorizontalPodAutoscaler autoscaling/v2 with metrics-server, scaling behavior and stabilization, VerticalPodAutoscaler modes, and in-place Pod resize.
Resources, QoS and quotas
Container requests and limits, CPU throttling and OOM kills, QoS classes and eviction order, LimitRange defaults and bounds, ResourceQuota on compute and object counts, and how admission rejects Pods.