Asterrr's Handbook

Troubleshooting networking

Tracing a failed request hop by hop, from DNS through the Service and its EndpointSlices to kube-proxy, NetworkPolicies and the CNI, with the test Pod commands that prove each hop.

Exam tasks: 5.5 (troubleshoot services and networking)

The decision: which hop drops the request: name resolution, the Service's selector and ports, the node's Service rules (kube-proxy), a NetworkPolicy, or Pod networking itself?

Follow the request

Test from inside the cluster, one hop at a time:

k run nettest --rm -it --image=busybox:1.37 --restart=Never -n shop -- sh
# inside the Pod:
nslookup web                          # 1: DNS (short name works only in the same namespace)
nslookup web.shop.svc.cluster.local
wget -qO- -T 3 http://web:8080        # 2-4: through the Service
wget -qO- -T 3 http://10.244.1.17:80  # 4: straight to a Pod IP, bypassing the Service
ResultHop to fix
Name doesn't resolve, Pod IP worksDNS (step 1)
Name resolves, Service times out, Pod IP worksService ports, endpoints or kube-proxy (2 to 4)
connection refused from the ServiceNo ready endpoint, or targetPort hits a closed port
Pod IP also times outNetworkPolicy or CNI (5, 6)

Services without endpoints

The most common break: the Service exists but routes to nothing.

k get svc web -n shop -o wide                                  # SELECTOR column
k get endpointslices -n shop -l kubernetes.io/service-name=web # ENDPOINTS empty?
k get pods -n shop --show-labels                               # do labels match the selector?
k get pods -n shop -l app=web -o wide                          # are matching Pods READY?
k describe svc web -n shop                                     # TargetPort, Endpoints
  • Selector mismatch. app: web in the Service, app: web-frontend on the Pods. Fix the Service selector (k edit svc) or the Pod template labels, whichever the task says is authoritative.
  • Pods not ready. Pods failing readiness are listed as not ready and get no traffic. Fix the probe or the app (Troubleshooting applications).
  • Wrong targetPort. Endpoints exist but connections are refused: targetPort must equal the port the container listens on. A named targetPort (http) must match a ports[].name in the Pod spec.
  • Wrong namespace. A Service only selects Pods in its own namespace.

Legacy: use EndpointSlices (discovery.k8s.io/v1) instead

The v1 Endpoints API is deprecated since v1.33. k get endpoints still works but prints a deprecation warning. Use k get endpointslices -l kubernetes.io/service-name=<svc>; k describe svc still shows the endpoint addresses either way.

port vs targetPort vs nodePort

port is what clients call on the ClusterIP, targetPort is the container's port, nodePort is opened on every node (30000 to 32767). A Service with port: 80, targetPort: 80 in front of an app on 8080 has endpoints and looks fine in k get svc, yet every request is refused. Check containerPort (or the app's actual listen port with k exec ... -- netstat -tlnp if available) against targetPort.

DNS failures

k -n kube-system get pods -l k8s-app=kube-dns -o wide    # CoreDNS Pods Running and Ready?
k -n kube-system logs -l k8s-app=kube-dns --tail=30      # plugin errors, loop detection, upstream timeouts
k -n kube-system get svc kube-dns                         # ClusterIP, often 10.96.0.10
k -n kube-system get cm coredns -o yaml                   # the Corefile
k exec <client-pod> -n shop -- cat /etc/resolv.conf       # nameserver = kube-dns IP, search domains, ndots:5
  • CoreDNS Pods down, CrashLoopBackOff on a Corefile typo, or scaled to 0: fix and restart with k -n kube-system rollout restart deploy coredns.
  • nameserver in /etc/resolv.conf comes from the kubelet's clusterDNS setting. If it doesn't match the kube-dns Service IP, every lookup fails on that node.
  • Cross-namespace names need at least svc.ns. A bare web only resolves inside shop.
  • Pod and Service name formats are on the CoreDNS page.

kube-proxy

kube-proxy runs as a DaemonSet and turns Services into packet rules on each node.

k -n kube-system get ds kube-proxy                    # DESIRED = READY?
k -n kube-system logs -l k8s-app=kube-proxy --tail=20
k -n kube-system get cm kube-proxy -o yaml | grep mode   # "" (iptables), "nftables" or "ipvs"
# on a node: are there rules for the Service?
sudo iptables-save | grep 'shop/web'                  # iptables mode
sudo nft list ruleset | grep -c 'shop/web'            # nftables mode
  • A kube-proxy Pod missing on one node breaks Services only from Pods on that node. Look for taints or a failed image pull on the DaemonSet's Pod there.
  • mode empty means iptables on Linux. nftables mode is GA since v1.33 and is the migration path away from ipvs.

Legacy: use the nftables (or iptables) proxy mode instead

kube-proxy's ipvs mode is deprecated as of v1.35. It still works but logs a warning at start-up.

NetworkPolicies and the CNI

k get netpol -A
k describe netpol -n shop              # podSelector, policyTypes, allowed peers and ports
  • A Pod selected by any policy with Ingress in policyTypes accepts only what some policy allows. Same for Egress.
  • An egress default-deny also blocks DNS. Allow UDP and TCP port 53 to the kube-dns Pods, or names stop resolving while IPs still look "allowed".
  • Policies only work if the CNI enforces them (Calico, Cilium and others do; plain flannel doesn't).
  • CNI failures show up earlier: Pods stuck ContainerCreating with "failed to setup network for sandbox", or nodes NotReady with "cni plugin not initialized". See Troubleshooting nodes and Network policies.

Exam signal

To prove a NetworkPolicy fix, test both what should work and what should still be blocked. Start the test Pod with the labels the policy expects, for example k run t --rm -it --image=busybox:1.37 --restart=Never -n shop -l role=frontend -- wget -qO- -T 3 web:8080.

Scenarios

Scenario
Service ledger in namespace books has type ClusterIP, port 80, targetPort 8080 and selector app=ledger. Pods labelled app=ledger,tier=api are Running and Ready and listen on 8080. `k get endpointslices -n books -l kubernetes.io/service-name=ledger` lists their IPs, yet requests from a Pod in namespace audit to http://ledger time out with 'bad address'. What's wrong?
Scenario
After a NetworkPolicy named lockdown was applied in namespace orders, Pods there can no longer reach api.payments.svc.cluster.local, but `wget` to the payments Service's ClusterIP still works. The policy has policyTypes: [Egress] and allows TCP 443 to namespace payments only. What should you add?
Scenario
Service cart in namespace web has no endpoints. Its selector is app=cart. The Deployment's Pods are labelled app=cart and show READY 0/1, with events 'Readiness probe failed: HTTP probe failed with statuscode: 404'. What restores traffic?

Drill

kubectl config use-context lab-net. In namespace keep, Service vault-web should send traffic on port 80 to the Pods of Deployment vault-web, which listen on port 9000. Requests to http://vault-web.keep currently fail. Fix the Service only.

Further reading

On this page