Cluster architecture
Control plane and node components, etcd and quorum, the API server as the hub, and the reconciliation loop that makes Kubernetes declarative.
Exam tasks: 1.1 (Kubernetes core concepts: cluster components and how they cooperate)
The decision: given a job described in the question (store state, pick a node, restart a container, route a Service), which component owns it?
The two halves of a cluster
- The control plane holds desired state and makes decisions. In production it runs on several dedicated machines; in a lab it can be one node, or Pods on that node.
- Nodes (worker nodes) run your Pods. Each one runs a kubelet, a container runtime and usually kube-proxy.
- Every arrow points at the API server. Components never talk to each other directly, and only the API server reads and writes etcd.
Control plane components
| Component | Its job | Not its job |
|---|---|---|
| kube-apiserver | Authenticates, authorizes and validates every request, runs admission, then persists objects in etcd. Serves watches so other components see changes | Deciding placement or running containers |
| etcd | Consistent, distributed key-value store for all cluster data (objects, not container data) | Storing application volumes or logs |
| kube-scheduler | Watches for Pods with no node, picks the best node, writes spec.nodeName | Starting the Pod. It only records a decision |
| kube-controller-manager | Runs the built-in controllers in one binary: Deployment, ReplicaSet, Job, Node, EndpointSlice, ServiceAccount and others | Placing Pods on nodes |
| cloud-controller-manager | Cloud-specific controllers: node lifecycle from the cloud API, routes, and cloud load balancers for type: LoadBalancer Services | Anything on bare metal clusters, where it's simply absent |
Exam signal
"Stores the state of the cluster" or "source of truth" is etcd. "The only component that talks to etcd" or "front end of the control plane" is the kube-apiserver. "Watches for unscheduled Pods" is the kube-scheduler.
Node components
- kubelet: the node agent. It registers the node, watches the API for Pods bound to it, asks the runtime to
start their containers, runs liveness, readiness and startup probes, restarts failed containers according to
restartPolicy, and reports status back. - Container runtime: pulls images and runs containers. The kubelet talks to it through the CRI. See Containers and runtimes.
- kube-proxy: programs iptables, IPVS or nftables rules so traffic to a Service's virtual IP reaches a backing Pod. Some CNI plugins (Cilium, for example) replace it entirely with eBPF.
- Add-ons run as ordinary Pods: CoreDNS for cluster DNS, the CNI plugin's agents, metrics-server.
kubelet vs kube-proxy
Both run on every node and both start with "kube", which makes them classic distractors. The kubelet manages Pods and containers. kube-proxy only handles Service networking. Neither one decides where Pods go.
- Static Pods are the one exception to "everything goes through the API server". The kubelet runs manifests
it finds in a local directory (usually
/etc/kubernetes/manifests) on its own. kubeadm runs the control plane itself this way, and the API only shows a read-only mirror Pod.
etcd and quorum
- etcd uses the Raft consensus algorithm. A write succeeds only when a majority (quorum) of members agree.
- Run an odd number of members: 3 tolerates one failure, 5 tolerates two. Going from 3 to 4 adds cost and no extra fault tolerance.
- Losing quorum makes the cluster read-only at best: running Pods keep running, but nothing new can be created, scaled or scheduled.
- Back etcd up. It's the only stateful control plane component; the others can be rebuilt from it.
The reconciliation loop
- Every controller runs the same loop: watch the objects it owns, compare
spec(desired) withstatus(observed), and act to close the gap. This is why Kubernetes is called declarative. - Controllers are level-triggered: they converge on the current desired state no matter how many events they missed, so a controller that restarts simply catches up.
- Controllers chain: a Deployment controller creates a ReplicaSet, the ReplicaSet controller creates Pods, the scheduler binds them, the kubelet runs them. Each step is a separate loop.
- Operators apply the same pattern to your own resource types through CustomResourceDefinitions (see CRDs and operators for depth).
Exam signal
"If a Pod managed by a ReplicaSet is deleted, a new one appears" is the reconciliation loop in action. The component doing it is the controller manager (ReplicaSet controller), not the kubelet and not the scheduler.
Scenarios
Three members need two for quorum, so one survivor can't commit writes: no new Pods, no scaling, no updates. Kubelets don't talk to etcd at all, and they keep already running containers alive. One member can't act alone, and the scheduler never stores data.
Placing Pods is the scheduler's only job; it writes the chosen node into spec.nodeName. The controller manager
creates the Pod objects (for example from a ReplicaSet) but leaves them unscheduled. The kubelet only acts on
Pods already bound to its node. The cloud controller manager handles cloud resources such as load balancers.
The cloud controller manager runs the service controller that provisions cloud load balancers, plus node and route controllers. kube-proxy only programs node-local rules for Service traffic. The API server stores the Service but doesn't call cloud APIs, and etcd is just storage.
Further reading
Domain 1 · Kubernetes fundamentals
44% of the exam. What each cluster component and object does, how you talk to the API, how the scheduler places Pods, and how containers and runtimes fit underneath.
Core objects
Pods, ReplicaSets, Deployments, StatefulSets, DaemonSets, Jobs and CronJobs, Services, ConfigMaps and Secrets, labels, annotations and namespaces, and how to tell look-alikes apart.