Asterrr's Handbook

Highly available control plane

Stacked vs external etcd, quorum and fault tolerance, the load-balanced control plane endpoint, joining extra control plane nodes with upload-certs and certificate keys, and leader election.

Exam tasks: 1.5 (implement and configure a highly available control plane)

The decision: how many control plane nodes, where etcd runs, and what single address the nodes and clients use so that losing one machine doesn't take the API down?

Two topologies

Stacked etcdExternal etcd
Machines for 3-way HA36 (3 control plane, 3 etcd)
etcd runs asA static Pod on each control plane node, talking to its local API serverIts own cluster, managed outside kubeadm's control plane join
Failure couplingLosing a node loses an API server and an etcd memberControl plane and etcd fail independently
Setup effortkubeadm init and join --control-plane handle everythingBuild etcd first, copy its certs, set etcd.external in the kubeadm config
kubeadm defaultYesNo

Quorum

etcd needs a majority of members (floor(n/2) + 1) to accept writes. Without quorum the API server can still serve some cached reads, but nothing can change.

MembersQuorumFailures tolerated
110
220
321
431
532
743

Exam signal

Use an odd number of members. A fourth member raises the quorum without adding fault tolerance, and an even split across two sites can leave neither side with a majority. Three tolerates one failure, five tolerates two.

The control plane endpoint

Every kubelet, kube-proxy and kubeconfig needs one stable address that reaches whichever API servers are alive.

  • Put a TCP load balancer (HAProxy with keepalived, kube-vip, or a cloud load balancer) in front of port 6443 on all control plane nodes, with health checks on /livez or a plain TCP check.
  • Pass that address as --control-plane-endpoint (or controlPlaneEndpoint in ClusterConfiguration) at init time. It goes into the API server certificate SANs and every generated kubeconfig.
  • Use a DNS name rather than a raw IP, so the load balancer can move later.

Adding HA to a cluster built without an endpoint

A cluster initialised without --control-plane-endpoint has kubeconfigs and certificates that point at the first node's own IP. kubeadm join --control-plane refuses to add a second control plane to it. Fixing that means reissuing the API server certificate and editing every kubeconfig, so set the endpoint from day one if HA is even a possibility.

Building it with kubeadm

# cp-1
sudo kubeadm init --control-plane-endpoint "k8s-api.corp.lab:6443" \
  --upload-certs --pod-network-cidr=10.40.0.0/16
  • --upload-certs encrypts the shared CA, service account and front-proxy keys and stores them in the kubeadm-certs Secret in kube-system. The output prints a certificate key that decrypts them.
  • The output gives two join commands: one with --control-plane --certificate-key ... for control plane nodes, and one without for workers.
# cp-2 and cp-3 (same node prep and packages as cp-1)
sudo kubeadm join k8s-api.corp.lab:6443 --token <token> \
  --discovery-token-ca-cert-hash sha256:<hash> \
  --control-plane --certificate-key <key>

The uploaded certificates are deleted after two hours. Later joins need new ones:

# on an existing control plane
sudo kubeadm init phase upload-certs --upload-certs        # prints a new certificate key
sudo kubeadm token create --print-join-command --certificate-key <new-key>

Without --upload-certs you copy /etc/kubernetes/pki/{ca,sa,front-proxy-ca}.* and pki/etcd/ca.* to the new node yourself before joining with --control-plane.

Treating the certificate key like a token

The certificate key decrypts the cluster CA private key. Anyone who has it within the two-hour window, plus a valid bootstrap token, can join a control plane node. Don't paste it into tickets, and regenerate it only when you're about to join.

Who's active: leader election

  • kube-apiserver is active-active. Every instance serves requests behind the load balancer.
  • kube-controller-manager and kube-scheduler are active-passive. Each instance competes for a Lease in kube-system; only the holder works, and the others take over if it stops renewing.
  • etcd elects its own Raft leader. Writes go through the leader; any member can serve linearizable reads by checking with it.
k get lease -n kube-system kube-scheduler kube-controller-manager \
  -o custom-columns=NAME:.metadata.name,HOLDER:.spec.holderIdentity
sudo etcdctl --endpoints=https://127.0.0.1:2379 --cacert=/etc/kubernetes/pki/etcd/ca.crt \
  --cert=/etc/kubernetes/pki/etcd/server.crt --key=/etc/kubernetes/pki/etcd/server.key \
  member list -w table
sudo etcdctl ... endpoint status --cluster -w table     # shows which member is leader

Removing a control plane node

  1. k drain cp-3 --ignore-daemonsets, then on cp-3 sudo kubeadm reset. For stacked etcd, reset also removes the local etcd member from the cluster.
  2. k delete node cp-3, and take cp-3 out of the load balancer pool.
  3. If the machine died and reset couldn't run, remove the dead member yourself with etcdctl member remove <ID> from a healthy node, or the cluster keeps counting it toward quorum.
3
Minimum control plane nodes for stacked HA that survives one failure.
floor(n/2)+1
etcd quorum size.
2 hours
Lifetime of certificates uploaded with --upload-certs.
Lease
Object kube-scheduler and kube-controller-manager use for leader election.

Scenarios

Scenario
A platform team runs a stacked-etcd cluster on 4 control plane nodes, two in each of two server rooms. A network fault isolates the rooms from each other. What happens to writes to the API?
Scenario
Three days after building an HA cluster with `kubeadm init --upload-certs`, you try to add a third control plane node using the saved join command with --certificate-key. The join fails while downloading certificates. What do you run first on an existing control plane node?

Drill

ssh ha-cp1 (first control plane of cluster ha-lab, built with --control-plane-endpoint ha-api.lab:6443)

Machine ha-cp3 is prepared (runtime, packages, node prep done). Join it as a control plane node, then report which node currently holds the kube-scheduler Lease in /opt/answers/scheduler-leader.txt on ha-cp1.

Further reading

On this page