Asterrr's Handbook

Workload autoscaling

Manual scaling, HorizontalPodAutoscaler autoscaling/v2 with metrics-server, scaling behavior and stabilization, VerticalPodAutoscaler modes, and in-place Pod resize.

Exam tasks: 2.3 (configure workload autoscaling)

The decision: should the workload get more Pods (horizontal), bigger Pods (vertical), or a one-off replica change, and what has to be in place for the autoscaler to see any metrics at all?

Which scaler

HPAVPAIn-place resize
Changesreplicas on the targetContainer requests (and limits proportionally)One Pod's container resources
Ships with KubernetesYesNo, install from the autoscaler projectYes, stable in v1.35
TargetsDeployment, StatefulSet, ReplicaSet, anything with a scale subresourceSameA Pod
Restarts PodsNoIn Recreate mode, yesUsually no

Manual scaling

k scale deploy/ingest --replicas=6 -n data
k scale sts/ledger-db --replicas=3 -n data             # StatefulSets scale one ordinal at a time
k scale deploy/ingest --current-replicas=6 --replicas=2 -n data   # only if it is currently 6
  • If an HPA targets the workload, it overrides your manual change on its next sync. Edit the HPA's minReplicas/maxReplicas instead.
  • kubectl apply of a manifest that sets replicas fights the HPA. Leave replicas out of manifests for autoscaled workloads.

HorizontalPodAutoscaler

k autoscale deploy/ingest -n data --min=2 --max=12 --cpu=65%
k get hpa -n data          # TARGETS shows current/target, e.g. 41%/65%
k describe hpa ingest -n data

--cpu accepts a percentage (65%, utilization of requests) or a quantity (500m, average value), and --memory works the same way. Older kubectl releases only had --cpu-percent=65.

apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: ingest
  namespace: data
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: ingest
  minReplicas: 2
  maxReplicas: 12
  metrics:
  - type: Resource
    resource:
      name: cpu
      target:
        type: Utilization        # percent of the Pods' CPU *requests*
        averageUtilization: 65
  - type: Resource
    resource:
      name: memory
      target:
        type: AverageValue       # absolute value per Pod
        averageValue: 400Mi
  behavior:
    scaleDown:
      stabilizationWindowSeconds: 120   # default 300
      policies:
      - type: Pods
        value: 2
        periodSeconds: 60               # remove at most 2 Pods per minute

How it decides:

  • desired = ceil(currentReplicas × currentMetric / targetMetric). At 4 Pods averaging 130% against a 65% target, it asks for 8.
  • With several metrics it computes each and takes the highest replica count.
  • It ignores changes within a 10% tolerance of the target, and scale-down uses the highest recommendation from the last 5 minutes (the stabilization window), so it shrinks slowly by design.
  • Other metric types: Pods and Object (custom metrics) and External need a metrics adapter such as Prometheus Adapter or KEDA. Only Resource works with metrics-server alone.

TARGETS shows unknown

kubectl get hpa showing <unknown>/65% means the HPA can't compute utilization. The usual causes, in order: metrics-server isn't installed or ready (k top pods fails too), the target Pods have no CPU request (utilization is a percentage of requests), or the Pods have only just started. A container with only a CPU limit gets a request equal to that limit, so that case works too. In a multi-container Pod, every container needs a CPU request for Utilization to work.

metrics-server

k apply -f https://github.com/kubernetes-sigs/metrics-server/releases/latest/download/components.yaml
k -n kube-system rollout status deploy/metrics-server
k top nodes

In lab clusters whose kubelets use self-signed serving certificates, metrics-server can't scrape them. Adding --kubelet-insecure-tls to the metrics-server container args gets it working there; it isn't a production fix.

Exam signal

For an HPA task, check three things before you leave it: k top pods returns numbers, the target's containers have CPU requests, and k get hpa shows a real percentage instead of unknown.

Legacy: use autoscaling/v2 instead

autoscaling/v2beta1 and autoscaling/v2beta2 are removed (since v1.25 and v1.26). autoscaling/v1 still serves but only supports targetCPUUtilizationPercentage. Write new HPAs as autoscaling/v2, which adds multiple metrics, memory and behavior.

VerticalPodAutoscaler

  • An add-on from kubernetes/autoscaler: CRDs plus a recommender, an updater and an admission controller.
  • It watches usage and sets requests (limits move proportionally) through updatePolicy.updateMode:
ModeBehaviour
OffOnly writes recommendations to the VPA status
InitialApplies recommendations when Pods are created, never touches running Pods
RecreateEvicts Pods whose requests are far off, so they come back with new values
InPlaceOrRecreateTries an in-place resize first, evicts only if that fails (newer VPA releases)
AutoDeprecated alias from older releases; write Recreate or InPlaceOrRecreate explicitly
  • Don't let a VPA and an HPA both act on CPU or memory for the same workload: VPA raises requests, which lowers utilization, which makes the HPA scale in. Pair VPA with an HPA on custom metrics, or use VPA in Off mode for sizing advice.

In-place Pod resize

Stable in v1.35: you can change CPU and memory of a running Pod's containers without recreating the Pod.

k patch pod ingest-0 -n data --subresource=resize -p \
  '{"spec":{"containers":[{"name":"ingest","resources":{"requests":{"cpu":"750m"},"limits":{"cpu":"1500m"}}}]}}'
k get pod ingest-0 -n data -o jsonpath='{.status.containerStatuses[0].resources}{"\n"}'
  • The change goes through the resize subresource (kubectl edit pod ... --subresource=resize also works). An ordinary update of a Pod's resources is rejected.
  • resizePolicy per resource chooses NotRequired (default, apply live) or RestartContainer (restart that container to apply, common for memory in JVM apps).
  • The QoS class can't change. A resize that would turn a Guaranteed Pod into Burstable is refused.
  • Pod conditions show progress: PodResizePending (reason Deferred or Infeasible when the node lacks room) and PodResizeInProgress.
  • Resizing a Pod owned by a Deployment is temporary: the next rollout uses the template values. Change the template for anything permanent.
15 s
Default HPA sync period in kube-controller-manager.
300 s
Default scale-down stabilization window.
10%
Default tolerance around the target before the HPA acts.
1
Default minReplicas when you leave it out.

Scenarios

Scenario
You created an HPA for Deployment `thumbs` with `kubectl autoscale deploy/thumbs --min=2 --max=8 --cpu=60%`. `kubectl top pods` works, but `kubectl get hpa` shows `<unknown>/60%` and the replica count never changes. What is the most likely cause?
Scenario
An HPA targets 50% CPU for Deployment `sorter`, currently at 3 replicas with average utilization 120%. The HPA has min 2 and max 6. What replica count will it request?
Scenario
A team wants Pods of Deployment `reports` to get right-sized CPU and memory requests based on observed usage, but they don't want any running Pod restarted by the tool. Which setup fits?

Drill

kubectl config use-context drill-w3. Deployment render in namespace media runs 1 replica and has no resource settings. metrics-server is installed.

  1. Give the render container a CPU request of 200m.
  2. Create an HPA named render that keeps between 2 and 7 replicas at 70% average CPU utilization, and scales down by at most 1 Pod per 30 seconds.

Further reading

On this page