Installing a cluster with kubeadm
Preparing nodes (swap, IP forwarding, containerd and the cgroup driver, pkgs.k8s.io packages), kubeadm init and its config file, installing a CNI, joining workers and regenerating join commands.
Exam tasks: 1.2 (prepare underlying infrastructure for installing a Kubernetes cluster), 1.3 (create and manage Kubernetes clusters using kubeadm)
The decision: what has to be true on every node before kubeadm init or kubeadm join succeeds, and which
flags or config fields shape the cluster you get?
The build order
Nodes stay NotReady and CoreDNS stays Pending until a CNI plugin is installed. That's expected, not a fault.
Preparing the node
| Requirement | Why | How |
|---|---|---|
| 2 CPUs, 2 GiB RAM on control plane nodes | Preflight checks fail below that | kubeadm init reports it; --ignore-preflight-errors only for labs |
Unique hostname, MAC and product_uuid | Cloned VMs collide and nodes overwrite each other | hostname, ip link, cat /sys/class/dmi/id/product_uuid |
| Swap off (unless you deliberately configure the kubelet for swap) | The kubelet refuses to start with swap on by default | swapoff -a and comment the swap line in /etc/fstab |
net.ipv4.ip_forward = 1 | Pod traffic is routed through the node | A file in /etc/sysctl.d/, then sysctl --system |
| Container runtime | The kubelet talks to it through the CRI | containerd or CRI-O, running and enabled |
| cgroup v2 | v1.35 kubelets fail to start on cgroup v1 by default | stat -fc %T /sys/fs/cgroup prints cgroup2fs |
| Open ports | Components must reach each other | 6443 API server, 2379–2380 etcd, 10250 kubelet, 10257 controller manager, 10259 scheduler, 30000–32767 NodePorts |
sudo swapoff -a && sudo sed -i '/ swap / s/^/#/' /etc/fstab
echo 'net.ipv4.ip_forward = 1' | sudo tee /etc/sysctl.d/k8s.conf
sudo sysctl --systemSome CNI plugins also need the overlay and br_netfilter kernel modules loaded; their install docs say so.
containerd and the cgroup driver
The kubelet and the runtime must agree on the cgroup driver. With systemd as init and cgroup v2, use systemd.
sudo mkdir -p /etc/containerd
containerd config default | sudo tee /etc/containerd/config.toml
# containerd 2.x: section [plugins.'io.containerd.cri.v1.runtime'.containerd.runtimes.runc.options]
# containerd 1.x: section [plugins."io.containerd.grpc.v1.cri".containerd.runtimes.runc.options]
sudo sed -i 's/SystemdCgroup = false/SystemdCgroup = true/' /etc/containerd/config.toml
sudo systemctl restart containerd && sudo systemctl enable containerd- kubeadm writes
cgroupDriver: systemdinto the kubelet config by default. - Since v1.34 the kubelet asks the runtime which driver it uses (the
RuntimeConfigCRI call) and follows it when the runtime supports that, as containerd 2.x and current CRI-O do. Older runtimes fall back to the kubelet's own setting, so a mismatch there means Pods that restart in a loop. - v1.35 is the last minor release that supports containerd 1.x. Plan for containerd 2.x.
Containers restart every few minutes after init
A cluster that comes up and then has control plane Pods restarting constantly is the classic cgroup driver
mismatch: containerd left on cgroupfs while the kubelet uses systemd. Set SystemdCgroup = true, restart
containerd, then restart the kubelet.
Packages from pkgs.k8s.io
Each minor version has its own repository. The URL contains the minor version, so an upgrade later means changing this line.
sudo apt-get update && sudo apt-get install -y apt-transport-https ca-certificates curl gpg
sudo mkdir -p -m 755 /etc/apt/keyrings
curl -fsSL https://pkgs.k8s.io/core:/stable:/v1.35/deb/Release.key \
| sudo gpg --dearmor -o /etc/apt/keyrings/kubernetes-apt-keyring.gpg
echo 'deb [signed-by=/etc/apt/keyrings/kubernetes-apt-keyring.gpg] https://pkgs.k8s.io/core:/stable:/v1.35/deb/ /' \
| sudo tee /etc/apt/sources.list.d/kubernetes.list
sudo apt-get update
sudo apt-get install -y kubelet kubeadm kubectl
sudo apt-mark hold kubelet kubeadm kubectl
sudo systemctl enable --now kubeletThe kubelet restarts every few seconds until kubeadm init or join gives it a config. That's normal.
Legacy: use pkgs.k8s.io instead
apt.kubernetes.io and yum.kubernetes.io (the Google-hosted repos) were frozen in 2023 and later removed.
Commands that add packages.cloud.google.com keys no longer work. The same goes for Docker Engine as a runtime
through dockershim, removed in v1.24: use containerd or CRI-O, or cri-dockerd if you must keep Docker.
kubeadm init
sudo kubeadm init \
--pod-network-cidr=10.32.0.0/16 \
--apiserver-advertise-address=192.168.56.10 \
--control-plane-endpoint=cp.lab.internal:6443 # only if you may add control plane nodes later
mkdir -p $HOME/.kube
sudo cp /etc/kubernetes/admin.conf $HOME/.kube/config
sudo chown $(id -u):$(id -g) $HOME/.kube/config| Flag | What it sets |
|---|---|
--pod-network-cidr | Pod address range given to nodes. Must match what your CNI expects |
--service-cidr | ClusterIP range. Default 10.96.0.0/12 |
--apiserver-advertise-address | IP the API server listens on and advertises. Pick the right NIC on multi-homed hosts |
--control-plane-endpoint | Shared address for all control plane nodes. Required for HA, hard to add later |
--kubernetes-version | Pin a version instead of the latest for this kubeadm |
--cri-socket | Which runtime, when more than one is installed |
--config | A kubeadm config file instead of flags |
For anything beyond a few flags, use a config file. The current API is kubeadm.k8s.io/v1beta4:
apiVersion: kubeadm.k8s.io/v1beta4
kind: ClusterConfiguration
kubernetesVersion: v1.35.0
controlPlaneEndpoint: cp.lab.internal:6443
networking:
podSubnet: 10.32.0.0/16
serviceSubnet: 10.96.0.0/12
apiServer:
extraArgs:
- name: audit-log-maxage # v1beta4 uses a list of name/value pairs
value: "14"
---
apiVersion: kubelet.config.k8s.io/v1beta1
kind: KubeletConfiguration
cgroupDriver: systemdkubeadm config print init-defaults prints a starting file, and sudo kubeadm init --config kubeadm.yaml uses it.
What init leaves behind
| Path | Contents |
|---|---|
/etc/kubernetes/manifests/ | Static Pod manifests for etcd, kube-apiserver, kube-controller-manager and kube-scheduler. The kubelet watches this folder |
/etc/kubernetes/pki/ | Cluster CA, API server, etcd and front-proxy certificates and keys |
/etc/kubernetes/admin.conf | Admin kubeconfig (bound to cluster-admin via the kubeadm:cluster-admins group) |
/etc/kubernetes/super-admin.conf | Break-glass kubeconfig in system:masters. Don't hand it out |
/var/lib/kubelet/config.yaml | The kubelet's configuration |
Installing a CNI
Install exactly one plugin, using its manifest, Helm chart or operator, and match its Pod CIDR to
--pod-network-cidr. Pick one that supports NetworkPolicy (Calico, Cilium) if the cluster will need policies.
k apply -f <manifest-from-the-plugin-docs>.yaml
k get pods -n kube-system -w # wait for the CNI Pods and CoreDNS to run
k get nodes # ReadyJoining nodes
# on the control plane: tokens from init expire after 24 hours
sudo kubeadm token create --print-join-command
# then, on the new worker (same node prep and packages as the control plane)
sudo kubeadm join cp.lab.internal:6443 --token 7x1kqp.2m9fz0rd8ycw3hjt \
--discovery-token-ca-cert-hash sha256:<hash-printed-by-the-command-above>kubeadm token listshows existing tokens and their TTL.- The CA hash pins the cluster CA, so a node can't be tricked into joining a fake control plane.
- Undo a failed attempt with
sudo kubeadm reseton that node, then clean/etc/cni/net.dand iptables or IPVS rules as the reset output tells you.
Exam signal
"Join node X to the cluster" almost always means the original token has expired. Run kubeadm token create --print-join-command from a control plane node, ssh to the worker, run the printed line with sudo, then come back
and check k get nodes.
Scenarios
Without a network plugin the kubelet reports the network as not ready, and CoreDNS can't get Pod IPs. That is the normal state between init and the CNI install. Resetting changes nothing, and switching the cgroup driver would create a real mismatch.
Bootstrap tokens from init live 24 hours. A new token with the same CA hash fixes it in one command. Copying the admin kubeconfig to a worker hands out cluster-admin, and init on the worker builds a second, separate cluster.
Drill
ssh lab-node2 (prepared host, packages installed, not yet in any cluster). The control plane is lab-cp1.
Join lab-node2 to the cluster as a worker. The original join command is no longer valid. Confirm the node is
Ready.
ssh lab-cp1
sudo kubeadm token create --print-join-command
# copy the printed line
exit
ssh lab-node2
systemctl is-active containerd # active
swapon --show # no output
sudo kubeadm join 192.168.56.10:6443 --token <new-token> \
--discovery-token-ca-cert-hash sha256:<hash>
exitVerify from the base host or the control plane:
k get nodes -o wide # lab-node2 Ready (allow a minute for the CNI Pod on it)
k get pods -n kube-system -o wide --field-selector spec.nodeName=lab-node2If the node stays NotReady, check sudo journalctl -u kubelet -e on lab-node2.
Further reading
RBAC
Roles and ClusterRoles, RoleBindings and ClusterRoleBindings, users, groups and ServiceAccounts as subjects, aggregated ClusterRoles, and proving access with kubectl auth can-i.
Cluster upgrades
Version skew rules, switching the pkgs.k8s.io repo, kubeadm upgrade plan, apply and node, draining and uncordoning, and renewing kubeadm certificates.