Asterrr's Handbook

Cluster architecture

Control plane and node components, etcd and quorum, the API server as the hub, and the reconciliation loop that makes Kubernetes declarative.

Exam tasks: 1.1 (Kubernetes core concepts: cluster components and how they cooperate)

The decision: given a job described in the question (store state, pick a node, restart a container, route a Service), which component owns it?

The two halves of a cluster

  • The control plane holds desired state and makes decisions. In production it runs on several dedicated machines; in a lab it can be one node, or Pods on that node.
  • Nodes (worker nodes) run your Pods. Each one runs a kubelet, a container runtime and usually kube-proxy.
  • Every arrow points at the API server. Components never talk to each other directly, and only the API server reads and writes etcd.

Control plane components

ComponentIts jobNot its job
kube-apiserverAuthenticates, authorizes and validates every request, runs admission, then persists objects in etcd. Serves watches so other components see changesDeciding placement or running containers
etcdConsistent, distributed key-value store for all cluster data (objects, not container data)Storing application volumes or logs
kube-schedulerWatches for Pods with no node, picks the best node, writes spec.nodeNameStarting the Pod. It only records a decision
kube-controller-managerRuns the built-in controllers in one binary: Deployment, ReplicaSet, Job, Node, EndpointSlice, ServiceAccount and othersPlacing Pods on nodes
cloud-controller-managerCloud-specific controllers: node lifecycle from the cloud API, routes, and cloud load balancers for type: LoadBalancer ServicesAnything on bare metal clusters, where it's simply absent

Exam signal

"Stores the state of the cluster" or "source of truth" is etcd. "The only component that talks to etcd" or "front end of the control plane" is the kube-apiserver. "Watches for unscheduled Pods" is the kube-scheduler.

Node components

  • kubelet: the node agent. It registers the node, watches the API for Pods bound to it, asks the runtime to start their containers, runs liveness, readiness and startup probes, restarts failed containers according to restartPolicy, and reports status back.
  • Container runtime: pulls images and runs containers. The kubelet talks to it through the CRI. See Containers and runtimes.
  • kube-proxy: programs iptables, IPVS or nftables rules so traffic to a Service's virtual IP reaches a backing Pod. Some CNI plugins (Cilium, for example) replace it entirely with eBPF.
  • Add-ons run as ordinary Pods: CoreDNS for cluster DNS, the CNI plugin's agents, metrics-server.

kubelet vs kube-proxy

Both run on every node and both start with "kube", which makes them classic distractors. The kubelet manages Pods and containers. kube-proxy only handles Service networking. Neither one decides where Pods go.

  • Static Pods are the one exception to "everything goes through the API server". The kubelet runs manifests it finds in a local directory (usually /etc/kubernetes/manifests) on its own. kubeadm runs the control plane itself this way, and the API only shows a read-only mirror Pod.

etcd and quorum

  • etcd uses the Raft consensus algorithm. A write succeeds only when a majority (quorum) of members agree.
  • Run an odd number of members: 3 tolerates one failure, 5 tolerates two. Going from 3 to 4 adds cost and no extra fault tolerance.
  • Losing quorum makes the cluster read-only at best: running Pods keep running, but nothing new can be created, scaled or scheduled.
  • Back etcd up. It's the only stateful control plane component; the others can be rebuilt from it.
floor(n/2) + 1
Members needed for etcd quorum. 3 members need 2, 5 need 3.
6443
Default secure port of the kube-apiserver.
2379
etcd client port that the API server connects to.
10250
kubelet API port, used by the API server for logs, exec and port-forward.

The reconciliation loop

  • Every controller runs the same loop: watch the objects it owns, compare spec (desired) with status (observed), and act to close the gap. This is why Kubernetes is called declarative.
  • Controllers are level-triggered: they converge on the current desired state no matter how many events they missed, so a controller that restarts simply catches up.
  • Controllers chain: a Deployment controller creates a ReplicaSet, the ReplicaSet controller creates Pods, the scheduler binds them, the kubelet runs them. Each step is a separate loop.
  • Operators apply the same pattern to your own resource types through CustomResourceDefinitions (see CRDs and operators for depth).

Exam signal

"If a Pod managed by a ReplicaSet is deleted, a new one appears" is the reconciliation loop in action. The component doing it is the controller manager (ReplicaSet controller), not the kubelet and not the scheduler.

Scenarios

Scenario
A platform team at Lumen Freight runs three etcd members. Two of the three machines lose power at the same time. Which statement describes the cluster?
Scenario
Which component watches the API for newly created Pods that have no node assigned and selects a node for each one?
Scenario
A team creates a Service of type LoadBalancer on a managed cloud cluster, and a cloud load balancer appears a minute later. Which component called the cloud provider's API to create it?

Further reading

On this page