etcd backup and restore
Finding etcd's endpoint and certificates, saving a snapshot with etcdctl, checking it and restoring it with etcdutl, and pointing the etcd static Pod at the restored data directory.
Exam tasks: 1.4 (manage the lifecycle of Kubernetes clusters)
The decision: where is etcd listening, which certificates does etcdctl need to talk to it, and what must
change in the static Pod manifest so the cluster runs from the restored data?
The two tools
| Task | Tool | Needs certificates? |
|---|---|---|
| Take a snapshot | etcdctl snapshot save | Yes: CA, client cert, client key, endpoint |
| Check a snapshot (hash, revision, keys, size) | etcdutl snapshot status | No |
| Restore to a new data directory | etcdutl snapshot restore | No |
| Member list, health | etcdctl member list, etcdctl endpoint health | Yes |
Legacy: use etcdutl snapshot restore and etcdutl snapshot status instead
etcdctl snapshot restore and etcdctl snapshot status were deprecated in etcd 3.5 and removed in etcd 3.6,
which is the etcd that kubeadm deploys for v1.35. Only snapshot save stays in etcdctl. You may also see
ETCDCTL_API=3 in older commands; v3 has been the default since etcd 3.4, so it's harmless but unnecessary.
Finding the endpoint and certificates
Never guess the paths: read them from the etcd static Pod manifest on the control plane node.
ssh cp-main
sudo grep -E 'listen-client-urls|cert-file|key-file|trusted-ca-file|data-dir' /etc/kubernetes/manifests/etcd.yaml| etcd flag in the manifest | etcdctl flag | Typical kubeadm value |
|---|---|---|
--listen-client-urls | --endpoints | https://127.0.0.1:2379 |
--trusted-ca-file | --cacert | /etc/kubernetes/pki/etcd/ca.crt |
--cert-file | --cert | /etc/kubernetes/pki/etcd/server.crt |
--key-file | --key | /etc/kubernetes/pki/etcd/server.key |
--data-dir | (restore target) | /var/lib/etcd |
The server certificate works as a client certificate on kubeadm clusters. So do
/etc/kubernetes/pki/etcd/healthcheck-client.crt and the API server's /etc/kubernetes/pki/apiserver-etcd-client.crt.
Saving a snapshot
sudo etcdctl --endpoints=https://127.0.0.1:2379 \
--cacert=/etc/kubernetes/pki/etcd/ca.crt \
--cert=/etc/kubernetes/pki/etcd/server.crt \
--key=/etc/kubernetes/pki/etcd/server.key \
snapshot save /srv/etcd/snap-0412.db
sudo etcdutl --write-out=table snapshot status /srv/etcd/snap-0412.dbsudomatters: the keys under/etc/kubernetes/pkiare readable by root only.- The snapshot comes from one member. On an HA cluster, any healthy member is fine.
- No
etcdctlon the host? The etcd image ships it. Run it in the Pod and write under/var/lib/etcd, which is a hostPath mount, so the file appears on the node:k -n kube-system exec etcd-cp-main -- etcdctl --endpoints=... --cacert=... --cert=... --key=... snapshot save /var/lib/etcd/snap.db
Exam signal
If snapshot save hangs or fails with context deadline exceeded, a TLS flag is missing or wrong. Rerun
etcdctl ... endpoint health with the same flags: if health fails too, check the cert paths against the manifest.
Restoring on a kubeadm control plane
ssh cp-main
sudo mkdir -p /root/mf-hold
sudo mv /etc/kubernetes/manifests/kube-apiserver.yaml /root/mf-hold/
sudo etcdutl snapshot restore /srv/etcd/snap-0412.db --data-dir /var/lib/etcd-restored
sudo vi /etc/kubernetes/manifests/etcd.yamlChange only the hostPath of the data volume. The container still sees /var/lib/etcd, so --data-dir and the
volumeMounts stay as they are:
volumes:
- hostPath:
path: /var/lib/etcd-restored # was /var/lib/etcd
type: DirectoryOrCreate
name: etcd-data# wait until etcd is running again
sudo crictl ps --name etcd
sudo mv /root/mf-hold/kube-apiserver.yaml /etc/kubernetes/manifests/
# give the API server a minute, then
k get nodes
k get deploy -A # objects from the snapshot time are back- The kubelet restarts etcd as soon as the manifest changes. Editing a copy elsewhere does nothing.
- Restoring into a new directory keeps the old data as a fallback and avoids "data dir not empty" errors.
- Restarting kube-controller-manager, kube-scheduler and the kubelets afterwards is recommended, so nothing works from stale cached state.
Restoring into the live data directory
Running the restore with --data-dir /var/lib/etcd while etcd is still running, or into a directory that already
has data, either fails or mixes old and new state. Restore into a fresh path and switch the hostPath. If you
change --data-dir in the command section instead, you must also change the volume mount, or etcd starts empty.
Multi-member restores
On a multi-member cluster, restore the same snapshot on every member, each with its own identity, so the members form a new cluster together:
sudo etcdutl snapshot restore /srv/etcd/snap-0412.db \
--name cp-main \
--initial-cluster cp-main=https://10.20.0.11:2380,cp-b=https://10.20.0.12:2380,cp-c=https://10.20.0.13:2380 \
--initial-advertise-peer-urls https://10.20.0.11:2380 \
--data-dir /var/lib/etcd-restoredTake --name and the URLs from each node's etcd.yaml. For a single-member lab cluster you can leave them out.
Scenarios
kubeadm's etcd requires TLS client authentication, so etcdctl needs the CA and a client certificate and key. API
version 3 is already the default, --data-dir is a restore-time flag, and 2380 is the peer port, not for clients.
The flag is the path inside the container. Only /var/lib/etcd is mounted from the host, so etcd created an empty
store at the new container path. Changing the hostPath (and leaving the flag alone) is the simple fix. Snapshots
contain every key, namespaced or not.
Drill
ssh cka-etcd-cp
- Save a snapshot of the cluster's etcd to
/opt/vault/etcd-pre.db. - Then restore the snapshot at
/opt/vault/etcd-golden.db(already on the node) so the cluster runs from it. Use/var/lib/etcd-goldenas the new data directory.
ssh cka-etcd-cp
sudo grep -E 'listen-client-urls|cert-file|key-file|trusted-ca-file' /etc/kubernetes/manifests/etcd.yaml
# 1. save
sudo etcdctl --endpoints=https://127.0.0.1:2379 \
--cacert=/etc/kubernetes/pki/etcd/ca.crt \
--cert=/etc/kubernetes/pki/etcd/server.crt \
--key=/etc/kubernetes/pki/etcd/server.key \
snapshot save /opt/vault/etcd-pre.db
sudo etcdutl --write-out=table snapshot status /opt/vault/etcd-pre.db
# 2. restore
sudo mkdir -p /root/mf-hold
sudo mv /etc/kubernetes/manifests/kube-apiserver.yaml /root/mf-hold/
sudo etcdutl snapshot restore /opt/vault/etcd-golden.db --data-dir /var/lib/etcd-golden
sudo sed -i 's#path: /var/lib/etcd$#path: /var/lib/etcd-golden#' /etc/kubernetes/manifests/etcd.yaml
sudo grep -A2 'hostPath' /etc/kubernetes/manifests/etcd.yaml # check only the etcd-data volume changed
sudo crictl ps --name etcd # wait until Running
sudo mv /root/mf-hold/kube-apiserver.yaml /etc/kubernetes/manifests/Verify:
k get nodes
k get all -A | head # objects match the golden snapshot, not the current state
sudo crictl ps | grep -E 'etcd|kube-apiserver'Further reading
Cluster upgrades
Version skew rules, switching the pkgs.k8s.io repo, kubeadm upgrade plan, apply and node, draining and uncordoning, and renewing kubeadm certificates.
Highly available control plane
Stacked vs external etcd, quorum and fault tolerance, the load-balanced control plane endpoint, joining extra control plane nodes with upload-certs and certificate keys, and leader election.