Highly available control plane
Stacked vs external etcd, quorum and fault tolerance, the load-balanced control plane endpoint, joining extra control plane nodes with upload-certs and certificate keys, and leader election.
Exam tasks: 1.5 (implement and configure a highly available control plane)
The decision: how many control plane nodes, where etcd runs, and what single address the nodes and clients use so that losing one machine doesn't take the API down?
Two topologies
| Stacked etcd | External etcd | |
|---|---|---|
| Machines for 3-way HA | 3 | 6 (3 control plane, 3 etcd) |
| etcd runs as | A static Pod on each control plane node, talking to its local API server | Its own cluster, managed outside kubeadm's control plane join |
| Failure coupling | Losing a node loses an API server and an etcd member | Control plane and etcd fail independently |
| Setup effort | kubeadm init and join --control-plane handle everything | Build etcd first, copy its certs, set etcd.external in the kubeadm config |
| kubeadm default | Yes | No |
Quorum
etcd needs a majority of members (floor(n/2) + 1) to accept writes. Without quorum the API server can still
serve some cached reads, but nothing can change.
| Members | Quorum | Failures tolerated |
|---|---|---|
| 1 | 1 | 0 |
| 2 | 2 | 0 |
| 3 | 2 | 1 |
| 4 | 3 | 1 |
| 5 | 3 | 2 |
| 7 | 4 | 3 |
Exam signal
Use an odd number of members. A fourth member raises the quorum without adding fault tolerance, and an even split across two sites can leave neither side with a majority. Three tolerates one failure, five tolerates two.
The control plane endpoint
Every kubelet, kube-proxy and kubeconfig needs one stable address that reaches whichever API servers are alive.
- Put a TCP load balancer (HAProxy with keepalived, kube-vip, or a cloud load balancer) in front of port 6443
on all control plane nodes, with health checks on
/livezor a plain TCP check. - Pass that address as
--control-plane-endpoint(orcontrolPlaneEndpointinClusterConfiguration) at init time. It goes into the API server certificate SANs and every generated kubeconfig. - Use a DNS name rather than a raw IP, so the load balancer can move later.
Adding HA to a cluster built without an endpoint
A cluster initialised without --control-plane-endpoint has kubeconfigs and certificates that point at the first
node's own IP. kubeadm join --control-plane refuses to add a second control plane to it. Fixing that means
reissuing the API server certificate and editing every kubeconfig, so set the endpoint from day one if HA is
even a possibility.
Building it with kubeadm
# cp-1
sudo kubeadm init --control-plane-endpoint "k8s-api.corp.lab:6443" \
--upload-certs --pod-network-cidr=10.40.0.0/16--upload-certsencrypts the shared CA, service account and front-proxy keys and stores them in thekubeadm-certsSecret inkube-system. The output prints a certificate key that decrypts them.- The output gives two join commands: one with
--control-plane --certificate-key ...for control plane nodes, and one without for workers.
# cp-2 and cp-3 (same node prep and packages as cp-1)
sudo kubeadm join k8s-api.corp.lab:6443 --token <token> \
--discovery-token-ca-cert-hash sha256:<hash> \
--control-plane --certificate-key <key>The uploaded certificates are deleted after two hours. Later joins need new ones:
# on an existing control plane
sudo kubeadm init phase upload-certs --upload-certs # prints a new certificate key
sudo kubeadm token create --print-join-command --certificate-key <new-key>Without --upload-certs you copy /etc/kubernetes/pki/{ca,sa,front-proxy-ca}.* and pki/etcd/ca.* to the new
node yourself before joining with --control-plane.
Treating the certificate key like a token
The certificate key decrypts the cluster CA private key. Anyone who has it within the two-hour window, plus a valid bootstrap token, can join a control plane node. Don't paste it into tickets, and regenerate it only when you're about to join.
Who's active: leader election
- kube-apiserver is active-active. Every instance serves requests behind the load balancer.
- kube-controller-manager and kube-scheduler are active-passive. Each instance competes for a Lease in
kube-system; only the holder works, and the others take over if it stops renewing. - etcd elects its own Raft leader. Writes go through the leader; any member can serve linearizable reads by checking with it.
k get lease -n kube-system kube-scheduler kube-controller-manager \
-o custom-columns=NAME:.metadata.name,HOLDER:.spec.holderIdentity
sudo etcdctl --endpoints=https://127.0.0.1:2379 --cacert=/etc/kubernetes/pki/etcd/ca.crt \
--cert=/etc/kubernetes/pki/etcd/server.crt --key=/etc/kubernetes/pki/etcd/server.key \
member list -w table
sudo etcdctl ... endpoint status --cluster -w table # shows which member is leaderRemoving a control plane node
k drain cp-3 --ignore-daemonsets, then on cp-3sudo kubeadm reset. For stacked etcd, reset also removes the local etcd member from the cluster.k delete node cp-3, and take cp-3 out of the load balancer pool.- If the machine died and reset couldn't run, remove the dead member yourself with
etcdctl member remove <ID>from a healthy node, or the cluster keeps counting it toward quorum.
Scenarios
Four members need three for quorum, and a 2–2 split leaves neither side with three, so all writes stop until the link returns. The previous leader steps down because it can't reach a majority. This is why odd member counts, and an odd number of failure domains, are recommended.
The uploaded certificates expired two hours after init. The upload-certs phase re-uploads them and prints a new
key; combine it with a fresh token from kubeadm token create --print-join-command. Renewing certificates doesn't
re-upload them, resetting destroys the control plane, and deleting the Secret makes things worse.
Drill
ssh ha-cp1 (first control plane of cluster ha-lab, built with --control-plane-endpoint ha-api.lab:6443)
Machine ha-cp3 is prepared (runtime, packages, node prep done). Join it as a control plane node, then
report which node currently holds the kube-scheduler Lease in /opt/answers/scheduler-leader.txt on ha-cp1.
ssh ha-cp1
KEY=$(sudo kubeadm init phase upload-certs --upload-certs | tail -1)
sudo kubeadm token create --print-join-command --certificate-key "$KEY"
# copy the printed line
exit
ssh ha-cp3
sudo kubeadm join ha-api.lab:6443 --token <token> \
--discovery-token-ca-cert-hash sha256:<hash> \
--control-plane --certificate-key <key>
exit
ssh ha-cp1
k get nodes # ha-cp3 Ready, role control-plane
k get pods -n kube-system -o wide | grep ha-cp3 # etcd, apiserver, controller-manager, scheduler
k get lease kube-scheduler -n kube-system -o jsonpath='{.spec.holderIdentity}' \
| sudo tee /opt/answers/scheduler-leader.txtThe holder identity looks like ha-cp1_<uuid>; the node name is the part before the underscore. Check etcd sees
three members with etcdctl member list using the certificate flags above.
Further reading
etcd backup and restore
Finding etcd's endpoint and certificates, saving a snapshot with etcdctl, checking it and restoring it with etcdutl, and pointing the etcd static Pod at the restored data directory.
Helm and Kustomize
Installing, upgrading and rolling back Helm releases with values, inspecting charts and releases, and building Kustomize bases and overlays you apply with kubectl apply -k.