Deployments, rolling updates and rollbacks
How a Deployment rolls out a new revision through ReplicaSets, tuning maxSurge and maxUnavailable, Recreate, kubectl rollout history, undo, pause and restart, and spotting a stalled rollout.
Exam tasks: 2.1 (Deployments, rolling updates and rollbacks)
The decision: how do you change a running app's version so that capacity and downtime stay within the limits the task sets, and how do you get back to a known-good revision when it goes wrong?
How a rollout works
- A revision is created only when the Pod template changes. Scaling, or editing
strategy, doesn't create one. - Each revision is a ReplicaSet named
<deployment>-<pod-template-hash>. Old ReplicaSets stay at 0 replicas sorollout undocan scale them back up.revisionHistoryLimit(default 10) caps how many are kept. - The
selectoris immutable inapps/v1. To change it, delete and recreate the Deployment.
Strategy parameters
| Field | Default | Effect |
|---|---|---|
strategy.type | RollingUpdate | Recreate kills every old Pod before starting new ones: brief outage, never two versions at once |
rollingUpdate.maxSurge | 25% (rounded up) | How many Pods above replicas may exist during the rollout |
rollingUpdate.maxUnavailable | 25% (rounded down) | How many Pods below replicas may be not Ready during the rollout |
minReadySeconds | 0 | A new Pod counts as available only after it has been Ready this long |
progressDeadlineSeconds | 600 | After this many seconds without progress the Deployment reports ProgressDeadlineExceeded |
revisionHistoryLimit | 10 | Old ReplicaSets kept for rollback |
spec:
replicas: 6
minReadySeconds: 10
strategy:
type: RollingUpdate
rollingUpdate:
maxSurge: 2 # up to 8 Pods in total
maxUnavailable: 0 # never fewer than 6 availablemaxSurgeandmaxUnavailablecan't both be 0.maxUnavailable: 0withmaxSurge: 1is the "never lose capacity" setting, at the cost of a slower rollout and one Pod's worth of spare node capacity.- Use
Recreatewhen two versions must never run together, for example an app that migrates a database schema on start or holds aReadWriteOncevolume that only one node can mount.
Exam signal
"Zero downtime" or "at no point fewer than N Pods" means maxUnavailable: 0. "No more than N extra Pods" sets
maxSurge. "The old and new versions must never serve traffic at the same time" means Recreate.
Commands you'll use
k create deploy ledger --image=registry.local/ledger:2.3 --replicas=4 -n pay
k set image deploy/ledger ledger=registry.local/ledger:2.4 -n pay # container=image
k annotate deploy/ledger kubernetes.io/change-cause="ledger 2.4: new fee rules" -n pay
k rollout status deploy/ledger -n pay --timeout=120s
k rollout history deploy/ledger -n pay
k rollout history deploy/ledger -n pay --revision=3 # Pod template of revision 3
k rollout undo deploy/ledger -n pay # back to the previous revision
k rollout undo deploy/ledger -n pay --to-revision=2
k rollout pause deploy/ledger -n pay # batch several edits into one revision
k rollout resume deploy/ledger -n pay
k rollout restart deploy/ledger -n pay # new Pods, same spec- In
set image, the part before=is the container name, not the Deployment name. Check it withk get deploy ledger -o jsonpath='{.spec.template.spec.containers[*].name}'. rollout restartworks by stamping akubectl.kubernetes.io/restartedAtannotation on the Pod template, so it does create a new revision. Use it to pick up changed ConfigMaps consumed as env vars.rollout undocopies the old template forward as a new revision number. The revision you rolled back to disappears from the history list under its old number.- While paused, template edits pile up without rolling.
resumerolls them out as one revision.
Legacy: use the kubernetes.io/change-cause annotation instead
The --record flag, which stored the command line in change-cause, is deprecated. Set the annotation yourself
with kubectl annotate after each change so rollout history shows something useful.
Waiting for an automatic rollback
When a rollout passes progressDeadlineSeconds, Kubernetes only marks the Deployment's Progressing condition
as False. It doesn't roll back on its own. The broken Pods sit in ImagePullBackOff or CrashLoopBackOff
until you run rollout undo or fix the template.
Reading a rollout's state
k get deploy ledger -n pay # READY, UP-TO-DATE, AVAILABLE
k get rs -n pay -l app=ledger # one RS per revision, which one has the replicas
k describe deploy ledger -n pay # Conditions and the scaling events per RS| Column | Means |
|---|---|
UP-TO-DATE | Pods built from the current template |
AVAILABLE | Pods Ready for at least minReadySeconds |
| Two ReplicaSets both with replicas | A rollout is in progress, paused, or stuck |
Blue/green and canary with plain Deployments
- Canary: run a second Deployment (
ledger-canary, 1 replica) whose Pods share the label the Service selects on. Traffic splits roughly by Pod count. - Blue/green: run
ledger-blueandledger-greenwith atracklabel and switch the Service selector from one to the other withk patch svc. Rollback is switching it back. - For weighted splits that don't depend on Pod counts, use Gateway API
HTTPRouteweights; see Gateway API.
Scenarios
maxUnavailable: 0 keeps 5 Pods available, and maxSurge: 1 lets exactly one extra Pod start at a time,
which fits the spare capacity. Recreate drops to zero. maxSurge: 0, maxUnavailable: 1 drops to 4. The
defaults would allow 2 extra Pods (25% of 5 rounded up) and 1 unavailable.
The bad image tag stalls the rollout, and after progressDeadlineSeconds the Deployment only reports
ProgressDeadlineExceeded. rollout undo restores the previous template. Deleting the new ReplicaSet just
makes the controller recreate it, a paused Deployment wouldn't have started the new Pod, and scaling the old
ReplicaSet down causes an outage.
While a Deployment is paused, template edits don't trigger a rollout, so resuming rolls all three out as one
revision. Three separate set commands create up to three revisions. Scaling to 0 causes an outage and still
creates revisions per edit. rollout restart adds even more revisions.
Drill
kubectl config use-context drill-w1. In namespace atlas:
- Create Deployment
tileswith 4 replicas ofnginx:1.27-alpine, container nametiles. - Configure it so a rollout never has fewer than 4 available Pods and adds at most 1 Pod at a time.
- Update the image to
nginx:1.28-alpine, recording the change causeupgrade to 1.28. - Roll back to the first revision and confirm the image.
kubectl create deploy names the container after the image (nginx), so generate the YAML and fix the name
and strategy before applying.
kubectl config use-context drill-w1
k create ns atlas --dry-run=client -o yaml | k apply -f -
k create deploy tiles --image=nginx:1.27-alpine --replicas=4 -n atlas --dry-run=client -o yaml > tiles.yamlEdit tiles.yaml so the spec contains:
spec:
replicas: 4
strategy:
type: RollingUpdate
rollingUpdate:
maxSurge: 1
maxUnavailable: 0
template:
spec:
containers:
- name: tiles # was nginx
image: nginx:1.27-alpinek apply -f tiles.yaml
k rollout status deploy/tiles -n atlas
k set image deploy/tiles tiles=nginx:1.28-alpine -n atlas
k annotate deploy/tiles kubernetes.io/change-cause="upgrade to 1.28" -n atlas
k rollout status deploy/tiles -n atlas
k rollout history deploy/tiles -n atlas # revisions 1 and 2
k rollout undo deploy/tiles --to-revision=1 -n atlas
k rollout status deploy/tiles -n atlasVerify:
k get deploy tiles -n atlas -o jsonpath='{.spec.template.spec.containers[0].image}{"\n"}' # nginx:1.27-alpine
k get deploy tiles -n atlas -o jsonpath='{.spec.strategy.rollingUpdate}{"\n"}'
k rollout history deploy/tiles -n atlas # now lists 2 and 3: revision 1 moved forward as 3Further reading
Domain 2 · Workloads and scheduling
15% of the exam. Rolling out and rolling back Deployments, injecting configuration, autoscaling, self-healing controllers, and controlling where Pods land and how much they may use.
ConfigMaps and Secrets
Creating ConfigMaps and Secrets from literals, files and env files, consuming them as environment variables or volumes, Secret types, immutability, and when an update actually reaches a running Pod.