Resource usage
Measuring what nodes, Pods and containers actually consume with metrics-server and kubectl top, telling usage apart from requests, and using events to explain evictions and pressure.
Exam tasks: 5.3 (monitor cluster and application resource usage)
The decision: do you need live usage (kubectl top, from metrics-server), reserved capacity
(requests and limits in describe node), or history of what happened (events)?
Three different numbers
| Question | Command | Source |
|---|---|---|
| What is this Pod using right now? | k top pod, k top pod --containers | metrics-server, scraped from each kubelet |
| How full is this node, really? | k top node | metrics-server |
| How much has been promised on this node? | k describe node → Allocated resources | Sum of Pod requests and limits in the API |
| Why was a Pod evicted, OOM-killed or not scheduled? | k events, k describe | Event objects |
| Long-term trends, dashboards, alerts | Not kubectl | Prometheus or another monitoring stack |
- metrics-server keeps only the latest sample in memory. It is not a monitoring system and has no history.
kubectl topmemory is the working set, the number the kubelet uses for eviction decisions. It's not the same as whatfreeshows on the node.- CPU is shown in millicores (
250mis a quarter of a core); memory inMi.
describe node is not usage
Allocated resources in k describe node adds up requests. A node can show 95% CPU requests while
k top node shows 10% actual use, or the reverse if Pods run without requests. A task that asks "which Pod
uses the most" wants kubectl top, not the requests table.
kubectl top
k top node # every node: CPU(cores) CPU% MEMORY(bytes) MEMORY%
k top pod -n checkout # Pods in one namespace
k top pod -A --sort-by=memory # whole cluster, biggest first
k top pod -l tier=batch -A --sort-by=cpu --no-headers | head -1
k top pod relay-7f9 -n checkout --containers # per container inside one PodExam signal
"Write the name of the Pod using the most CPU among Pods labelled x=y to /opt/answers/top.txt" is a classic.
Use --sort-by=cpu --no-headers and take the first line. With -A, the first column is the namespace and the
second is the Pod name, so awk '{print $2}'; without -A it's $1. Write only what was asked.
When kubectl top doesn't work
| Error | Fix |
|---|---|
error: Metrics API not available | metrics-server isn't installed or isn't ready. Check k get apiservice v1beta1.metrics.k8s.io and k -n kube-system get deploy metrics-server |
APIService Available=False, "FailedDiscoveryCheck" | The metrics-server Pod isn't running or the API server can't reach it. Check its logs |
| metrics-server logs "x509: cannot validate certificate … doesn't contain any IP SANs" | Kubelet serving certs are self-signed. In a lab, add --kubelet-insecure-tls to its args; in production, sign kubelet serving certs |
metrics not available yet for a new Pod | Wait for the next scrape (default resolution 15s) |
# install the current release
k apply -f https://github.com/kubernetes-sigs/metrics-server/releases/latest/download/components.yaml
# lab-only: trust kubelets' self-signed serving certs
k -n kube-system patch deploy metrics-server --type=json \
-p='[{"op":"add","path":"/spec/template/spec/containers/0/args/-","value":"--kubelet-insecure-tls"}]'
k get --raw /apis/metrics.k8s.io/v1beta1/nodes | head -c 300 # raw API checkEvents: what happened and why
Events record scheduling failures, image pulls, probe failures, OOM kills, evictions and node pressure.
k events -n checkout # sorted oldest to newest
k events -n checkout --types=Warning # only problems
k events -n checkout --for pod/relay-7f9 # one object
k events -A --watch # live stream
k get events -n checkout --field-selector type=Warning,reason=Evicted
k get events -A --sort-by=.metadata.creationTimestamp- The API server keeps events for 1 hour by default (
--event-ttl). An incident from yesterday won't show. Evictedwith "The node was low on resource: memory" points to nodeMemoryPressure. Look at the node's conditions and the top memory consumers on it.OOMKilledis a container hitting its own memory limit, not the node running out. See Troubleshooting applications.
kube-apiserver --event-ttl).kubectl top and the HPA.kubectl top shows and the kubelet evicts on.Finding the noisy Pod on a node
k top node # which node is hot
k get pods -A -o wide --field-selector spec.nodeName=worker-4
k top pod -A --sort-by=memory | head # then match names to that node
k describe node worker-4 | grep -A6 Conditions # pressure flagsFor QoS and eviction order (BestEffort before Burstable before Guaranteed), see Resources and quotas.
Scenarios
The task asks about current use, which is what metrics-server reports through kubectl top. Requests and
limits are what Pods asked for or are capped at; neither says how much they use now. Restarts are unrelated.
metrics-server can't verify the kubelets' self-signed serving certificates. In a lab, telling it to skip verification fixes it at once (in production you would issue proper kubelet serving certificates instead). Restarting kubelets doesn't change their certs, a different tag has the same check, and kubectl top has no fallback without the Metrics API.
Drill
kubectl config use-context lab-metrics. Among all Pods in all namespaces with label team=fulcrum, find the
one using the most CPU and write only its name to /opt/drill/cpu-hog.txt.
k top pod -A -l team=fulcrum --sort-by=cpu
k top pod -A -l team=fulcrum --sort-by=cpu --no-headers | head -1 | awk '{print $2}' > /opt/drill/cpu-hog.txt
cat /opt/drill/cpu-hog.txtIf kubectl top fails with "Metrics API not available", check
k -n kube-system get deploy metrics-server and k get apiservice v1beta1.metrics.k8s.io, fix metrics-server
(see the table above), wait about 30 seconds, then rerun. Verify the file holds one line with a Pod name and no
namespace column.
Further reading
Troubleshooting the control plane
Finding and fixing a broken kube-apiserver, kube-scheduler, kube-controller-manager or etcd on a kubeadm cluster, through static Pod manifests, crictl, Pod log files and health endpoints.
Container logs
Reading container stdout and stderr with kubectl logs (previous instances, multi-container Pods, selectors, time windows), where the logs live on the node, rotation, and streaming file logs through a sidecar.