Pod networking
The Kubernetes networking model, how containers in a Pod share one network namespace, how the CNI plugin gives Pods routable IPs across nodes, and what kube-proxy adds on top.
Exam tasks: 3.1 (connectivity between Pods), with groundwork for 3.3 and 5.5
The decision: when a Pod can't reach another Pod, is the problem inside the Pod (ports, localhost), in the Pod network the CNI plugin builds, or in the Service layer that kube-proxy programs?
The model in four rules
Kubernetes doesn't ship a Pod network. It sets the rules, and a CNI plugin implements them:
- Every Pod gets its own IP, unique across the cluster.
- Any Pod can reach any other Pod, on any node, without NAT. The source IP the receiver sees is the sender's Pod IP.
- Agents on a node (the kubelet, for probes) can reach every Pod on that node.
- Containers in the same Pod share that one IP and talk over
localhost.
Inside one Pod
- The Pod's network namespace is created first and held by the sandbox (pause) container. App containers
join it, so they share the IP, the port space, the routing table and
localhost. - Two containers in one Pod can't both bind port 8080. The second one fails with "address already in use".
containerPortis informational. A process listening on a port you didn't declare still receives traffic. Declaring it matters for named ports that Services and probes reference.hostNetwork: trueskips the Pod namespace and puts the Pod on the node's IP. You'll see it on CNI agents, kube-proxy and static control-plane Pods. It also makes the Pod compete for host ports.
Exam signal
"The sidecar must scrape the app without going through a Service" means the sidecar calls localhost:<port>.
"Two containers in the Pod keep crashing, one with bind errors" means they're fighting over the same port in the
shared namespace.
Across nodes: the CNI plugin
| Piece | Where you find it |
|---|---|
| Plugin config | /etc/cni/net.d/ on each node (the lowest-sorted file wins) |
| Plugin binaries | /opt/cni/bin/ on each node |
| Agent | Usually a DaemonSet in kube-system or its own namespace (Calico, Cilium, Flannel...) |
| Pod CIDR for the cluster | kubeadm init --pod-network-cidr, must match what the plugin is configured with |
| Per-node range | kubectl get node <n> -o jsonpath='{.spec.podCIDR}' when the controller manager allocates them |
- The container runtime calls the CNI plugin when the kubelet creates a Pod sandbox. The plugin creates the interface, assigns an IP (IPAM) and sets routes.
- Plugins differ in how packets cross nodes (overlay, BGP routing, eBPF) and in features. Flannel gives you connectivity only; Calico and Cilium also enforce NetworkPolicy.
- Until a plugin is installed, nodes stay
NotReadywith a message such asnetwork plugin not ready/cni plugin not initialized, and CoreDNS stays Pending.
Pod CIDR overlapping something else
The Pod CIDR, the Service CIDR (--service-cluster-ip-range, 10.96.0.0/12 by default with kubeadm) and the
node network must not overlap. A plugin manifest left on its default range while kubeadm was given a different
--pod-network-cidr produces Pods that get IPs but can't reach each other across nodes.
kube-proxy: the Service layer, not the Pod layer
Pod-to-Pod traffic doesn't touch kube-proxy. kube-proxy only turns Service virtual IPs into Pod IPs:
- Runs as a DaemonSet in
kube-system, configured by thekube-proxyConfigMap (mode:field). - Watches Services and EndpointSlices and writes packet rules on each node.
- Some CNI plugins (Cilium in kube-proxy replacement mode) do this job themselves, so you may find no kube-proxy at all.
| Mode | Status on v1.35 | Notes |
|---|---|---|
iptables | Default on Linux | Rule chains per Service; the one you'll meet most |
nftables | GA, opt-in | Successor to iptables mode, scales better on large clusters |
ipvs | Deprecated (v1.35) | Kernel load balancer; plan to move to nftables |
Legacy: use nftables mode instead
The ipvs proxy mode is deprecated as of v1.35. Older material recommends it for large clusters; the current
recommendation for that case is nftables. Set the mode explicitly in the kube-proxy config so an upgrade
doesn't change it under you.
Checking connectivity fast
k get pods -o wide -n shop # Pod IPs and the node each runs on
k get nodes -o custom-columns=NAME:.metadata.name,CIDR:.spec.podCIDR
k exec -n shop checkout -c metrics-agent -- wget -qO- -T 2 localhost:8080/healthz
k run probe --rm -it --restart=Never --image=busybox:1.36 -- wget -qO- -T 2 http://10.42.2.9:8080
k -n kube-system get ds # is the CNI agent and kube-proxy running everywhere?Scenarios
Containers in a Pod share the sandbox's network namespace, so localhost reaches the api directly. Pod names
aren't DNS records on their own (only Services, and Pods behind a headless Service, get names). Nothing is
exposed on the node IP without hostNetwork, hostPort or a NodePort. Containers don't get separate IPs.
Nodes stay NotReady until a CNI plugin writes its config to /etc/cni/net.d, and CoreDNS waits on that. kube-proxy
handles Service IPs, not Pod networking. Tolerating the taint would schedule CoreDNS onto a node that still can't
give it an IP. The Service CIDR must be a different range from the Pod CIDR.
Direct Pod-IP traffic between nodes is the CNI plugin's job (routes, overlay tunnel, eBPF). kube-proxy only matters for Service IPs, and DNS isn't involved when you connect by IP. A wrong targetPort would break the Service on both nodes, not just cross-node Pod traffic.
Drill
kubectl config use-context lab-net
In namespace relay, create a Pod named twin with two containers: web (image nginx:1.27) and poller
(image busybox:1.36) that runs wget -qO- localhost:80 every 5 seconds and writes to stdout. Then write the
IP of twin and the name of the node it runs on, separated by a space, to /opt/answers/twin.txt.
apiVersion: v1
kind: Pod
metadata:
name: twin
namespace: relay
spec:
containers:
- name: web
image: nginx:1.27
ports:
- containerPort: 80
- name: poller
image: busybox:1.36
command: ["sh", "-c", "while true; do wget -qO- localhost:80 | head -n 4; sleep 5; done"]k create ns relay --dry-run=client -o yaml | k apply -f -
k apply -f twin.yaml
k -n relay wait --for=condition=Ready pod/twin --timeout=60s
k -n relay get pod twin -o jsonpath='{.status.podIP} {.spec.nodeName}{"\n"}' > /opt/answers/twin.txt
# verify
k -n relay logs twin -c poller --tail=4 # HTML from nginx means localhost works
cat /opt/answers/twin.txtFurther reading
Domain 3 · Services and networking
20% of the exam. How Pods reach each other, how traffic gets in through Services, Gateway API and Ingress, how NetworkPolicies fence it, and how CoreDNS turns names into IPs.
Network policies
Default-deny NetworkPolicies, selecting peers by Pod labels, namespace labels and CIDRs, the AND vs OR trap in from and to, and keeping DNS working when you lock down egress.