Workload autoscaling
Manual scaling, HorizontalPodAutoscaler autoscaling/v2 with metrics-server, scaling behavior and stabilization, VerticalPodAutoscaler modes, and in-place Pod resize.
Exam tasks: 2.3 (configure workload autoscaling)
The decision: should the workload get more Pods (horizontal), bigger Pods (vertical), or a one-off replica change, and what has to be in place for the autoscaler to see any metrics at all?
Which scaler
| HPA | VPA | In-place resize | |
|---|---|---|---|
| Changes | replicas on the target | Container requests (and limits proportionally) | One Pod's container resources |
| Ships with Kubernetes | Yes | No, install from the autoscaler project | Yes, stable in v1.35 |
| Targets | Deployment, StatefulSet, ReplicaSet, anything with a scale subresource | Same | A Pod |
| Restarts Pods | No | In Recreate mode, yes | Usually no |
Manual scaling
k scale deploy/ingest --replicas=6 -n data
k scale sts/ledger-db --replicas=3 -n data # StatefulSets scale one ordinal at a time
k scale deploy/ingest --current-replicas=6 --replicas=2 -n data # only if it is currently 6- If an HPA targets the workload, it overrides your manual change on its next sync. Edit the HPA's
minReplicas/maxReplicasinstead. kubectl applyof a manifest that setsreplicasfights the HPA. Leavereplicasout of manifests for autoscaled workloads.
HorizontalPodAutoscaler
k autoscale deploy/ingest -n data --min=2 --max=12 --cpu=65%
k get hpa -n data # TARGETS shows current/target, e.g. 41%/65%
k describe hpa ingest -n data--cpu accepts a percentage (65%, utilization of requests) or a quantity (500m, average value), and
--memory works the same way. Older kubectl releases only had --cpu-percent=65.
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: ingest
namespace: data
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: ingest
minReplicas: 2
maxReplicas: 12
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization # percent of the Pods' CPU *requests*
averageUtilization: 65
- type: Resource
resource:
name: memory
target:
type: AverageValue # absolute value per Pod
averageValue: 400Mi
behavior:
scaleDown:
stabilizationWindowSeconds: 120 # default 300
policies:
- type: Pods
value: 2
periodSeconds: 60 # remove at most 2 Pods per minuteHow it decides:
desired = ceil(currentReplicas × currentMetric / targetMetric). At 4 Pods averaging 130% against a 65% target, it asks for 8.- With several metrics it computes each and takes the highest replica count.
- It ignores changes within a 10% tolerance of the target, and scale-down uses the highest recommendation from the last 5 minutes (the stabilization window), so it shrinks slowly by design.
- Other metric types:
PodsandObject(custom metrics) andExternalneed a metrics adapter such as Prometheus Adapter or KEDA. OnlyResourceworks with metrics-server alone.
TARGETS shows unknown
kubectl get hpa showing <unknown>/65% means the HPA can't compute utilization. The usual causes, in order:
metrics-server isn't installed or ready (k top pods fails too), the target Pods have no CPU request
(utilization is a percentage of requests), or the Pods have only just started. A container with only a CPU
limit gets a request equal to that limit, so that case works too. In a multi-container Pod, every container
needs a CPU request for Utilization to work.
metrics-server
k apply -f https://github.com/kubernetes-sigs/metrics-server/releases/latest/download/components.yaml
k -n kube-system rollout status deploy/metrics-server
k top nodesIn lab clusters whose kubelets use self-signed serving certificates, metrics-server can't scrape them. Adding
--kubelet-insecure-tls to the metrics-server container args gets it working there; it isn't a production fix.
Exam signal
For an HPA task, check three things before you leave it: k top pods returns numbers, the target's containers
have CPU requests, and k get hpa shows a real percentage instead of unknown.
Legacy: use autoscaling/v2 instead
autoscaling/v2beta1 and autoscaling/v2beta2 are removed (since v1.25 and v1.26). autoscaling/v1 still
serves but only supports targetCPUUtilizationPercentage. Write new HPAs as autoscaling/v2, which adds
multiple metrics, memory and behavior.
VerticalPodAutoscaler
- An add-on from
kubernetes/autoscaler: CRDs plus a recommender, an updater and an admission controller. - It watches usage and sets requests (limits move proportionally) through
updatePolicy.updateMode:
| Mode | Behaviour |
|---|---|
Off | Only writes recommendations to the VPA status |
Initial | Applies recommendations when Pods are created, never touches running Pods |
Recreate | Evicts Pods whose requests are far off, so they come back with new values |
InPlaceOrRecreate | Tries an in-place resize first, evicts only if that fails (newer VPA releases) |
Auto | Deprecated alias from older releases; write Recreate or InPlaceOrRecreate explicitly |
- Don't let a VPA and an HPA both act on CPU or memory for the same workload: VPA raises requests, which
lowers utilization, which makes the HPA scale in. Pair VPA with an HPA on custom metrics, or use VPA in
Offmode for sizing advice.
In-place Pod resize
Stable in v1.35: you can change CPU and memory of a running Pod's containers without recreating the Pod.
k patch pod ingest-0 -n data --subresource=resize -p \
'{"spec":{"containers":[{"name":"ingest","resources":{"requests":{"cpu":"750m"},"limits":{"cpu":"1500m"}}}]}}'
k get pod ingest-0 -n data -o jsonpath='{.status.containerStatuses[0].resources}{"\n"}'- The change goes through the
resizesubresource (kubectl edit pod ... --subresource=resizealso works). An ordinary update of a Pod's resources is rejected. resizePolicyper resource choosesNotRequired(default, apply live) orRestartContainer(restart that container to apply, common for memory in JVM apps).- The QoS class can't change. A resize that would turn a Guaranteed Pod into Burstable is refused.
- Pod conditions show progress:
PodResizePending(reasonDeferredorInfeasiblewhen the node lacks room) andPodResizeInProgress. - Resizing a Pod owned by a Deployment is temporary: the next rollout uses the template values. Change the template for anything permanent.
Scenarios
Utilization is a percentage of the CPU request, so without requests there is nothing to compare against.
Because kubectl top works, metrics-server is fine, which rules out the TLS flag. v2 supports CPU percentages,
and minReplicas has no effect on whether metrics are read.
ceil(3 × 120 / 50) = ceil(7.2) = 8, but maxReplicas caps it at 6. 7 and 8 ignore the cap, and 3 would mean the HPA isn't acting even though utilization is far outside the 10% tolerance.
Initial applies recommendations only when new Pods are created, for example during the next rollout, and
never evicts running ones. Recreate evicts Pods to apply new values. An HPA changes the number of Pods, not
their size. A ResourceQuota caps totals and doesn't size anything.
Drill
kubectl config use-context drill-w3. Deployment render in namespace media runs 1 replica and has no
resource settings. metrics-server is installed.
- Give the
rendercontainer a CPU request of200m. - Create an HPA named
renderthat keeps between 2 and 7 replicas at 70% average CPU utilization, and scales down by at most 1 Pod per 30 seconds.
kubectl config use-context drill-w3
k set resources deploy/render -n media -c render --requests=cpu=200m
k rollout status deploy/render -n mediaapiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: render
namespace: media
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: render
minReplicas: 2
maxReplicas: 7
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70
behavior:
scaleDown:
policies:
- type: Pods
value: 1
periodSeconds: 30k apply -f render-hpa.yaml
k get hpa render -n media # TARGETS shows a number like 3%/70% after a minute
k get deploy render -n media # replicas raised to 2 by the HPA minimumkubectl autoscale deploy/render -n media --min=2 --max=7 --cpu=70% creates the same HPA without the
behavior block; add it afterwards with k edit hpa render -n media.
Further reading
ConfigMaps and Secrets
Creating ConfigMaps and Secrets from literals, files and env files, consuming them as environment variables or volumes, Secret types, immutability, and when an update actually reaches a running Pod.
Self-healing workloads
Choosing between ReplicaSet, Deployment, DaemonSet, StatefulSet, Job and CronJob, restart policies and backoff, liveness, readiness and startup probes, native sidecar containers, and PodDisruptionBudgets.