Resources, QoS and quotas
Container requests and limits, CPU throttling and OOM kills, QoS classes and eviction order, LimitRange defaults and bounds, ResourceQuota on compute and object counts, and how admission rejects Pods.
Exam tasks: 2.5 (configure Pod admission and scheduling: limits)
The decision: how much CPU and memory should each container reserve and be capped at, and which namespace guardrail (LimitRange or ResourceQuota) enforces that for everyone?
Requests, limits and where each one acts
| Request | Limit | |
|---|---|---|
| Used by | The scheduler: a node must have that much unreserved allocatable capacity | The kubelet and runtime, through cgroups |
| CPU | Guaranteed share under contention | Hard cap: the container is throttled, never killed |
| Memory | Counts toward node fit and eviction ranking | Hard cap: the container is OOMKilled (exit code 137) |
| Missing | Defaults to the limit if a limit is set, otherwise 0 | No cap |
resources:
requests:
cpu: 250m # 0.25 core
memory: 192Mi # Mi = 2^20 bytes; M = 10^6
limits:
cpu: "1"
memory: 384Mik set resources deploy/orbit -n space -c orbit --requests=cpu=250m,memory=192Mi --limits=cpu=1,memory=384Mi
k describe node worker-2 | grep -A8 'Allocated resources' # requests and limits already committed- A Pod that fits no node by requests stays
PendingwithInsufficient cpuorInsufficient memoryin its events, even if the nodes are idle in practice. Scheduling never looks at actual usage. cpu: 1andcpu: 1000mare equal.memory: 1Gis about 7% less than1Gi.
QoS classes
| Class | Rule | Under node memory pressure |
|---|---|---|
| Guaranteed | Every container has CPU and memory requests equal to limits | Evicted last |
| Burstable | At least one container has a CPU or memory request or limit, but not Guaranteed | Evicted by how far usage exceeds requests |
| BestEffort | No container has any request or limit | Evicted first |
k get pod orbit-7d9c -n space -o jsonpath='{.status.qosClass}{"\n"}'- QoS is computed from the spec at creation and can't change afterwards, not even by in-place resize.
- Setting only limits gives requests equal to limits, so a Pod whose every container sets only CPU and memory limits is Guaranteed.
Exam signal
"This Pod must be the last to be evicted" or "must get Guaranteed QoS" means requests equal limits for both CPU and memory, in every container including init and sidecar containers.
LimitRange: per-object defaults and bounds
apiVersion: v1
kind: LimitRange
metadata:
name: container-bounds
namespace: space
spec:
limits:
- type: Container
defaultRequest: # injected when a container sets no request
cpu: 100m
memory: 128Mi
default: # injected when a container sets no limit
cpu: 500m
memory: 256Mi
min:
cpu: 50m
memory: 64Mi
max:
cpu: "2"
memory: 1Gi
- type: PersistentVolumeClaim
max:
storage: 20Gi- Applied by the
LimitRangeradmission plugin when a Pod is created. Existing Pods are untouched, so restart the workload to see the defaults. - A container that asks for more than
max(or less thanmin) is rejected with aForbiddenerror. type: Podbounds the sum across containers;maxLimitRequestRatiocaps limit divided by request.
Default limit below the request
If a LimitRange sets default.memory: 256Mi and a container sets only requests.memory: 512Mi, the injected
limit is smaller than the request and the Pod is rejected. Either set an explicit limit on the container, or
make the LimitRange default at least as large as any request you expect.
ResourceQuota: namespace totals
k create quota space-quota -n space \
--hard=requests.cpu=4,requests.memory=8Gi,limits.cpu=8,limits.memory=16Gi,pods=30,services.loadbalancers=0
k describe quota space-quota -n space # Used vs Hard| Quota key | Caps |
|---|---|
requests.cpu, requests.memory, limits.cpu, limits.memory | Sums across all non-terminal Pods |
pods, services, configmaps, secrets, persistentvolumeclaims | Object counts |
count/deployments.apps, count/jobs.batch | Count of any namespaced resource, count/<resource>.<group> |
requests.storage, <class>.storageclass.storage.k8s.io/requests.storage | Total PVC storage, overall or per StorageClass |
services.loadbalancers, services.nodeports | Service types |
- With a quota on
requests.cpu, every new Pod must set a CPU request (onlimits.memory, a memory limit, and so on), or admission rejects it. A LimitRange with defaults is the usual companion. scopeSelectorlimits a quota to some Pods:BestEffort,NotBestEffort,Terminating,NotTerminating, or aPriorityClass(for example "at most 2 Pods with prioritycritical-batchin this namespace").- Quotas are checked at admission only. Lowering a quota below current use doesn't evict anything; it blocks new objects until usage drops.
Looking for the quota error on the Deployment
When a Deployment's Pods exceed the quota, kubectl get deploy just shows fewer Ready replicas. The
exceeded quota error is an event on the ReplicaSet (FailedCreate). Run
k describe rs -l app=orbit -n space or k get events -n space to see it.
The admission chain in short
- Requests pass authentication, authorization, then mutating admission (LimitRanger defaults, ServiceAccount, mutating webhooks), schema validation, then validating admission (ResourceQuota, Pod Security Admission, ValidatingAdmissionPolicy, validating webhooks). Any one can reject the object.
- Plugins are toggled on kube-apiserver with
--enable-admission-plugins/--disable-admission-plugins.LimitRangerandResourceQuotaare on by default. - Pod Security Admission is driven by namespace labels such as
pod-security.kubernetes.io/enforce=restricted; a Pod that breaks the level is rejected at creation.
Scenarios
A compute quota requires every new Pod to declare requests for the quota-tracked resources. Defaults from a LimitRange, or explicit requests, satisfy it. Deleting and recreating the quota leaves it unenforced for running Pods, a bigger quota doesn't help Pods that declare nothing, and priority doesn't affect quota admission.
Guaranteed needs CPU and memory requests equal to limits in every container. agent gets a memory request
equal to its limit, but has no CPU request or limit, so the Pod is Burstable. It isn't BestEffort because
resources are set, and no admission rule rejects a missing CPU request unless a quota or LimitRange demands one.
OOMKilled with exit code 137 means the container went over its memory limit. CPU over the limit is throttled, not killed, so raising CPU doesn't help. A probe only adds more restarts, and a lower request changes scheduling, not the cap.
Drill
kubectl config use-context drill-w5. In namespace quarry:
- Make sure any container created without resources gets a request of
100mCPU and128Mimemory and a limit of300mCPU and256Mimemory. - Limit the namespace to 10 Pods and a total of
2CPU in requests. - Create Pod
probe(imagenginx:1.27-alpine, no resources in the spec) and confirm what it received.
apiVersion: v1
kind: LimitRange
metadata:
name: quarry-defaults
namespace: quarry
spec:
limits:
- type: Container
defaultRequest:
cpu: 100m
memory: 128Mi
default:
cpu: 300m
memory: 256Mikubectl config use-context drill-w5
k apply -f quarry-limits.yaml
k create quota quarry-quota -n quarry --hard=pods=10,requests.cpu=2
k run probe -n quarry --image=nginx:1.27-alpine
k get pod probe -n quarry -o jsonpath='{.spec.containers[0].resources}{"\n"}'
# {"limits":{"cpu":"300m","memory":"256Mi"},"requests":{"cpu":"100m","memory":"128Mi"}}
k describe quota quarry-quota -n quarry # pods 1/10, requests.cpu 100m/2Create the LimitRange before the Pod. LimitRange defaults are injected only at creation.
Further reading
Self-healing workloads
Choosing between ReplicaSet, Deployment, DaemonSet, StatefulSet, Job and CronJob, restart policies and backoff, liveness, readiness and startup probes, native sidecar containers, and PodDisruptionBudgets.
Pod placement
Steering Pods with nodeSelector, node affinity and Pod anti-affinity, keeping them off nodes with taints and tolerations, spreading them with topology spread constraints, and ordering them with PriorityClass and preemption.