Asterrr's Handbook

etcd backup and restore

Finding etcd's endpoint and certificates, saving a snapshot with etcdctl, checking it and restoring it with etcdutl, and pointing the etcd static Pod at the restored data directory.

Exam tasks: 1.4 (manage the lifecycle of Kubernetes clusters)

The decision: where is etcd listening, which certificates does etcdctl need to talk to it, and what must change in the static Pod manifest so the cluster runs from the restored data?

The two tools

TaskToolNeeds certificates?
Take a snapshotetcdctl snapshot saveYes: CA, client cert, client key, endpoint
Check a snapshot (hash, revision, keys, size)etcdutl snapshot statusNo
Restore to a new data directoryetcdutl snapshot restoreNo
Member list, healthetcdctl member list, etcdctl endpoint healthYes

Legacy: use etcdutl snapshot restore and etcdutl snapshot status instead

etcdctl snapshot restore and etcdctl snapshot status were deprecated in etcd 3.5 and removed in etcd 3.6, which is the etcd that kubeadm deploys for v1.35. Only snapshot save stays in etcdctl. You may also see ETCDCTL_API=3 in older commands; v3 has been the default since etcd 3.4, so it's harmless but unnecessary.

Finding the endpoint and certificates

Never guess the paths: read them from the etcd static Pod manifest on the control plane node.

ssh cp-main
sudo grep -E 'listen-client-urls|cert-file|key-file|trusted-ca-file|data-dir' /etc/kubernetes/manifests/etcd.yaml
etcd flag in the manifestetcdctl flagTypical kubeadm value
--listen-client-urls--endpointshttps://127.0.0.1:2379
--trusted-ca-file--cacert/etc/kubernetes/pki/etcd/ca.crt
--cert-file--cert/etc/kubernetes/pki/etcd/server.crt
--key-file--key/etc/kubernetes/pki/etcd/server.key
--data-dir(restore target)/var/lib/etcd

The server certificate works as a client certificate on kubeadm clusters. So do /etc/kubernetes/pki/etcd/healthcheck-client.crt and the API server's /etc/kubernetes/pki/apiserver-etcd-client.crt.

Saving a snapshot

sudo etcdctl --endpoints=https://127.0.0.1:2379 \
  --cacert=/etc/kubernetes/pki/etcd/ca.crt \
  --cert=/etc/kubernetes/pki/etcd/server.crt \
  --key=/etc/kubernetes/pki/etcd/server.key \
  snapshot save /srv/etcd/snap-0412.db

sudo etcdutl --write-out=table snapshot status /srv/etcd/snap-0412.db
  • sudo matters: the keys under /etc/kubernetes/pki are readable by root only.
  • The snapshot comes from one member. On an HA cluster, any healthy member is fine.
  • No etcdctl on the host? The etcd image ships it. Run it in the Pod and write under /var/lib/etcd, which is a hostPath mount, so the file appears on the node: k -n kube-system exec etcd-cp-main -- etcdctl --endpoints=... --cacert=... --cert=... --key=... snapshot save /var/lib/etcd/snap.db

Exam signal

If snapshot save hangs or fails with context deadline exceeded, a TLS flag is missing or wrong. Rerun etcdctl ... endpoint health with the same flags: if health fails too, check the cert paths against the manifest.

Restoring on a kubeadm control plane

ssh cp-main
sudo mkdir -p /root/mf-hold
sudo mv /etc/kubernetes/manifests/kube-apiserver.yaml /root/mf-hold/

sudo etcdutl snapshot restore /srv/etcd/snap-0412.db --data-dir /var/lib/etcd-restored

sudo vi /etc/kubernetes/manifests/etcd.yaml

Change only the hostPath of the data volume. The container still sees /var/lib/etcd, so --data-dir and the volumeMounts stay as they are:

  volumes:
  - hostPath:
      path: /var/lib/etcd-restored      # was /var/lib/etcd
      type: DirectoryOrCreate
    name: etcd-data
# wait until etcd is running again
sudo crictl ps --name etcd
sudo mv /root/mf-hold/kube-apiserver.yaml /etc/kubernetes/manifests/
# give the API server a minute, then
k get nodes
k get deploy -A           # objects from the snapshot time are back
  • The kubelet restarts etcd as soon as the manifest changes. Editing a copy elsewhere does nothing.
  • Restoring into a new directory keeps the old data as a fallback and avoids "data dir not empty" errors.
  • Restarting kube-controller-manager, kube-scheduler and the kubelets afterwards is recommended, so nothing works from stale cached state.

Restoring into the live data directory

Running the restore with --data-dir /var/lib/etcd while etcd is still running, or into a directory that already has data, either fails or mixes old and new state. Restore into a fresh path and switch the hostPath. If you change --data-dir in the command section instead, you must also change the volume mount, or etcd starts empty.

Multi-member restores

On a multi-member cluster, restore the same snapshot on every member, each with its own identity, so the members form a new cluster together:

sudo etcdutl snapshot restore /srv/etcd/snap-0412.db \
  --name cp-main \
  --initial-cluster cp-main=https://10.20.0.11:2380,cp-b=https://10.20.0.12:2380,cp-c=https://10.20.0.13:2380 \
  --initial-advertise-peer-urls https://10.20.0.11:2380 \
  --data-dir /var/lib/etcd-restored

Take --name and the URLs from each node's etcd.yaml. For a single-member lab cluster you can leave them out.

2379
etcd client port. Peers talk on 2380.
3.6
etcd minor version kubeadm deploys for Kubernetes v1.35, where etcdctl lost restore and status.
/var/lib/etcd
Default kubeadm etcd data directory, mounted into the Pod by hostPath.

Scenarios

Scenario
You run `etcdctl --endpoints=https://127.0.0.1:2379 snapshot save /srv/snap.db` on a kubeadm control plane and it fails after a few seconds with context deadline exceeded. etcd is healthy. What is missing?
Scenario
After restoring a snapshot with `etcdutl snapshot restore snap.db --data-dir /var/lib/etcd-new`, you edit etcd.yaml and change the `--data-dir` flag to /var/lib/etcd-new. etcd starts, but the cluster has no Deployments at all, not even ones that existed before the snapshot. Why?

Drill

ssh cka-etcd-cp

  1. Save a snapshot of the cluster's etcd to /opt/vault/etcd-pre.db.
  2. Then restore the snapshot at /opt/vault/etcd-golden.db (already on the node) so the cluster runs from it. Use /var/lib/etcd-golden as the new data directory.

Further reading

On this page