Troubleshooting networking
Tracing a failed request hop by hop, from DNS through the Service and its EndpointSlices to kube-proxy, NetworkPolicies and the CNI, with the test Pod commands that prove each hop.
Exam tasks: 5.5 (troubleshoot services and networking)
The decision: which hop drops the request: name resolution, the Service's selector and ports, the node's Service rules (kube-proxy), a NetworkPolicy, or Pod networking itself?
Follow the request
Test from inside the cluster, one hop at a time:
k run nettest --rm -it --image=busybox:1.37 --restart=Never -n shop -- sh
# inside the Pod:
nslookup web # 1: DNS (short name works only in the same namespace)
nslookup web.shop.svc.cluster.local
wget -qO- -T 3 http://web:8080 # 2-4: through the Service
wget -qO- -T 3 http://10.244.1.17:80 # 4: straight to a Pod IP, bypassing the Service| Result | Hop to fix |
|---|---|
| Name doesn't resolve, Pod IP works | DNS (step 1) |
| Name resolves, Service times out, Pod IP works | Service ports, endpoints or kube-proxy (2 to 4) |
connection refused from the Service | No ready endpoint, or targetPort hits a closed port |
| Pod IP also times out | NetworkPolicy or CNI (5, 6) |
Services without endpoints
The most common break: the Service exists but routes to nothing.
k get svc web -n shop -o wide # SELECTOR column
k get endpointslices -n shop -l kubernetes.io/service-name=web # ENDPOINTS empty?
k get pods -n shop --show-labels # do labels match the selector?
k get pods -n shop -l app=web -o wide # are matching Pods READY?
k describe svc web -n shop # TargetPort, Endpoints- Selector mismatch.
app: webin the Service,app: web-frontendon the Pods. Fix the Service selector (k edit svc) or the Pod template labels, whichever the task says is authoritative. - Pods not ready. Pods failing readiness are listed as not ready and get no traffic. Fix the probe or the app (Troubleshooting applications).
- Wrong
targetPort. Endpoints exist but connections are refused:targetPortmust equal the port the container listens on. A namedtargetPort(http) must match aports[].namein the Pod spec. - Wrong namespace. A Service only selects Pods in its own namespace.
Legacy: use EndpointSlices (discovery.k8s.io/v1) instead
The v1 Endpoints API is deprecated since v1.33. k get endpoints still works but prints a deprecation
warning. Use k get endpointslices -l kubernetes.io/service-name=<svc>; k describe svc still shows the
endpoint addresses either way.
port vs targetPort vs nodePort
port is what clients call on the ClusterIP, targetPort is the container's port, nodePort is opened on
every node (30000 to 32767). A Service with port: 80, targetPort: 80 in front of an app on 8080 has
endpoints and looks fine in k get svc, yet every request is refused. Check containerPort (or the app's actual
listen port with k exec ... -- netstat -tlnp if available) against targetPort.
DNS failures
k -n kube-system get pods -l k8s-app=kube-dns -o wide # CoreDNS Pods Running and Ready?
k -n kube-system logs -l k8s-app=kube-dns --tail=30 # plugin errors, loop detection, upstream timeouts
k -n kube-system get svc kube-dns # ClusterIP, often 10.96.0.10
k -n kube-system get cm coredns -o yaml # the Corefile
k exec <client-pod> -n shop -- cat /etc/resolv.conf # nameserver = kube-dns IP, search domains, ndots:5- CoreDNS Pods down,
CrashLoopBackOffon a Corefile typo, or scaled to 0: fix and restart withk -n kube-system rollout restart deploy coredns. nameserverin/etc/resolv.confcomes from the kubelet'sclusterDNSsetting. If it doesn't match thekube-dnsService IP, every lookup fails on that node.- Cross-namespace names need at least
svc.ns. A barewebonly resolves insideshop. - Pod and Service name formats are on the CoreDNS page.
kube-proxy
kube-proxy runs as a DaemonSet and turns Services into packet rules on each node.
k -n kube-system get ds kube-proxy # DESIRED = READY?
k -n kube-system logs -l k8s-app=kube-proxy --tail=20
k -n kube-system get cm kube-proxy -o yaml | grep mode # "" (iptables), "nftables" or "ipvs"
# on a node: are there rules for the Service?
sudo iptables-save | grep 'shop/web' # iptables mode
sudo nft list ruleset | grep -c 'shop/web' # nftables mode- A kube-proxy Pod missing on one node breaks Services only from Pods on that node. Look for taints or a failed image pull on the DaemonSet's Pod there.
modeempty means iptables on Linux. nftables mode is GA since v1.33 and is the migration path away from ipvs.
Legacy: use the nftables (or iptables) proxy mode instead
kube-proxy's ipvs mode is deprecated as of v1.35. It still works but logs a warning at start-up.
NetworkPolicies and the CNI
k get netpol -A
k describe netpol -n shop # podSelector, policyTypes, allowed peers and ports- A Pod selected by any policy with
IngressinpolicyTypesaccepts only what some policy allows. Same forEgress. - An egress default-deny also blocks DNS. Allow UDP and TCP port 53 to the
kube-dnsPods, or names stop resolving while IPs still look "allowed". - Policies only work if the CNI enforces them (Calico, Cilium and others do; plain flannel doesn't).
- CNI failures show up earlier: Pods stuck
ContainerCreatingwith "failed to setup network for sandbox", or nodesNotReadywith "cni plugin not initialized". See Troubleshooting nodes and Network policies.
Exam signal
To prove a NetworkPolicy fix, test both what should work and what should still be blocked. Start the test
Pod with the labels the policy expects, for example
k run t --rm -it --image=busybox:1.37 --restart=Never -n shop -l role=frontend -- wget -qO- -T 3 web:8080.
Scenarios
'bad address' is a DNS failure. The short name ledger only resolves inside books, because the client's search
path starts with its own namespace. The selector already matches (extra Pod labels are fine), endpoints exist,
and port 80 to targetPort 8080 is a normal mapping. kube-proxy problems show up after resolution, as timeouts
or refused connections.
The egress policy now denies everything not listed, including DNS lookups to CoreDNS in kube-system, so names fail while the IP works. Allowing port 53 to the DNS Pods fixes it without opening anything else. Ingress rules govern traffic into orders, a second DNS Service changes nothing, and hostNetwork bypasses the problem by weakening isolation.
The selector matches, but Pods that fail readiness are not ready endpoints, so the Service has nowhere to send traffic. Fixing the probe path or the endpoint the app serves makes them Ready and populates the EndpointSlice. Broadening the selector, changing the Service type or restarting kube-proxy doesn't change Pod readiness.
Drill
kubectl config use-context lab-net. In namespace keep, Service vault-web should send traffic on port 80 to
the Pods of Deployment vault-web, which listen on port 9000. Requests to http://vault-web.keep currently
fail. Fix the Service only.
k get svc vault-web -n keep -o wide
k get endpointslices -n keep -l kubernetes.io/service-name=vault-web # ENDPOINTS <unset>
k get pods -n keep --show-labels
# example finding: Service selector app=vaultweb, Pods labelled app=vault-web; targetPort 8080
k get deploy vault-web -n keep -o jsonpath='{.spec.template.spec.containers[0].ports}'Patch the selector and the target port:
k patch svc vault-web -n keep --type=merge \
-p '{"spec":{"selector":{"app":"vault-web"}}}'
k patch svc vault-web -n keep --type=json \
-p '[{"op":"replace","path":"/spec/ports/0/targetPort","value":9000}]'A merge patch merges map keys, so app gets the new value and other selector keys stay. To drop a stale key,
set it to null in the patch ({"spec":{"selector":{"old":null}}}).
Verify:
k get endpointslices -n keep -l kubernetes.io/service-name=vault-web # Pod IPs on port 9000
k run t --rm -it --image=busybox:1.37 --restart=Never -n default -- wget -qO- -T 3 http://vault-web.keep