Asterrr's Handbook

Pod networking

The Kubernetes networking model, how containers in a Pod share one network namespace, how the CNI plugin gives Pods routable IPs across nodes, and what kube-proxy adds on top.

Exam tasks: 3.1 (connectivity between Pods), with groundwork for 3.3 and 5.5

The decision: when a Pod can't reach another Pod, is the problem inside the Pod (ports, localhost), in the Pod network the CNI plugin builds, or in the Service layer that kube-proxy programs?

The model in four rules

Kubernetes doesn't ship a Pod network. It sets the rules, and a CNI plugin implements them:

  1. Every Pod gets its own IP, unique across the cluster.
  2. Any Pod can reach any other Pod, on any node, without NAT. The source IP the receiver sees is the sender's Pod IP.
  3. Agents on a node (the kubelet, for probes) can reach every Pod on that node.
  4. Containers in the same Pod share that one IP and talk over localhost.

Inside one Pod

  • The Pod's network namespace is created first and held by the sandbox (pause) container. App containers join it, so they share the IP, the port space, the routing table and localhost.
  • Two containers in one Pod can't both bind port 8080. The second one fails with "address already in use".
  • containerPort is informational. A process listening on a port you didn't declare still receives traffic. Declaring it matters for named ports that Services and probes reference.
  • hostNetwork: true skips the Pod namespace and puts the Pod on the node's IP. You'll see it on CNI agents, kube-proxy and static control-plane Pods. It also makes the Pod compete for host ports.

Exam signal

"The sidecar must scrape the app without going through a Service" means the sidecar calls localhost:<port>. "Two containers in the Pod keep crashing, one with bind errors" means they're fighting over the same port in the shared namespace.

Across nodes: the CNI plugin

PieceWhere you find it
Plugin config/etc/cni/net.d/ on each node (the lowest-sorted file wins)
Plugin binaries/opt/cni/bin/ on each node
AgentUsually a DaemonSet in kube-system or its own namespace (Calico, Cilium, Flannel...)
Pod CIDR for the clusterkubeadm init --pod-network-cidr, must match what the plugin is configured with
Per-node rangekubectl get node <n> -o jsonpath='{.spec.podCIDR}' when the controller manager allocates them
  • The container runtime calls the CNI plugin when the kubelet creates a Pod sandbox. The plugin creates the interface, assigns an IP (IPAM) and sets routes.
  • Plugins differ in how packets cross nodes (overlay, BGP routing, eBPF) and in features. Flannel gives you connectivity only; Calico and Cilium also enforce NetworkPolicy.
  • Until a plugin is installed, nodes stay NotReady with a message such as network plugin not ready / cni plugin not initialized, and CoreDNS stays Pending.

Pod CIDR overlapping something else

The Pod CIDR, the Service CIDR (--service-cluster-ip-range, 10.96.0.0/12 by default with kubeadm) and the node network must not overlap. A plugin manifest left on its default range while kubeadm was given a different --pod-network-cidr produces Pods that get IPs but can't reach each other across nodes.

kube-proxy: the Service layer, not the Pod layer

Pod-to-Pod traffic doesn't touch kube-proxy. kube-proxy only turns Service virtual IPs into Pod IPs:

  • Runs as a DaemonSet in kube-system, configured by the kube-proxy ConfigMap (mode: field).
  • Watches Services and EndpointSlices and writes packet rules on each node.
  • Some CNI plugins (Cilium in kube-proxy replacement mode) do this job themselves, so you may find no kube-proxy at all.
ModeStatus on v1.35Notes
iptablesDefault on LinuxRule chains per Service; the one you'll meet most
nftablesGA, opt-inSuccessor to iptables mode, scales better on large clusters
ipvsDeprecated (v1.35)Kernel load balancer; plan to move to nftables

Legacy: use nftables mode instead

The ipvs proxy mode is deprecated as of v1.35. Older material recommends it for large clusters; the current recommendation for that case is nftables. Set the mode explicitly in the kube-proxy config so an upgrade doesn't change it under you.

Checking connectivity fast

k get pods -o wide -n shop                 # Pod IPs and the node each runs on
k get nodes -o custom-columns=NAME:.metadata.name,CIDR:.spec.podCIDR
k exec -n shop checkout -c metrics-agent -- wget -qO- -T 2 localhost:8080/healthz
k run probe --rm -it --restart=Never --image=busybox:1.36 -- wget -qO- -T 2 http://10.42.2.9:8080
k -n kube-system get ds                    # is the CNI agent and kube-proxy running everywhere?
1 IP
Per Pod, shared by all its containers.
No NAT
Between Pods, on the same or different nodes.
/etc/cni/net.d
Where the runtime looks for CNI configuration.
/opt/cni/bin
Where CNI plugin binaries live.
10.96.0.0/12
kubeadm default Service CIDR. The Pod CIDR has no default; you pass it.

Scenarios

Scenario
Pod ledger runs two containers: api listening on 9000 and a log-shipper that must push buffered records to the api over HTTP. There is no Service for ledger. Which address should log-shipper use?
Scenario
After building a kubeadm cluster with --pod-network-cidr=10.244.0.0/16, both worker nodes report NotReady and the CoreDNS Pods are Pending. kubectl describe node shows 'container runtime network not ready'. What should you do?
Scenario
A Pod on worker-1 can reach Pods on worker-1 but times out reaching any Pod on worker-2, by IP. Services that select only worker-1 Pods work. Which component is the FIRST to check?

Drill

kubectl config use-context lab-net

In namespace relay, create a Pod named twin with two containers: web (image nginx:1.27) and poller (image busybox:1.36) that runs wget -qO- localhost:80 every 5 seconds and writes to stdout. Then write the IP of twin and the name of the node it runs on, separated by a space, to /opt/answers/twin.txt.

Further reading

On this page