Asterrr's Handbook

CoreDNS

DNS names for Services and Pods, how a Pod's resolv.conf is built, dnsPolicy and dnsConfig, the CoreDNS Corefile and its plugins, forwarding a private zone, and a step-by-step for debugging resolution.

Exam tasks: 3.6 (use CoreDNS), with 5.5 when names don't resolve

The decision: which name should the client use, which resolver answers it (CoreDNS's kubernetes plugin, a forwarded upstream, or the node's resolver), and where do you change that?

What's running

PieceName in a kubeadm cluster
Deploymentcoredns in kube-system (Pods labelled k8s-app=kube-dns)
Servicekube-dns in kube-system. Its ClusterIP (often 10.96.0.10) is what Pods use as nameserver
ConfigConfigMap coredns, key Corefile
Kubelet sideclusterDNS and clusterDomain in /var/lib/kubelet/config.yaml

Names you can resolve

RecordNameAnswer
Servicepricing.sales.svc.cluster.localThe ClusterIP
Headless Serviceledger-peers.sales.svc.cluster.localOne A record per ready Pod
StatefulSet Pod behind headless Serviceledger-0.ledger-peers.sales.svc.cluster.localThat Pod's IP
Named port (SRV)_http._tcp.pricing.sales.svc.cluster.localPort number and target
Any Pod by IP10-42-1-14.sales.pod.cluster.local10.42.1.14
ExternalName Servicebilling-api.sales.svc.cluster.localCNAME to the external name
  • A Pod gets the <hostname>.<subdomain>.<ns>.svc.cluster.local name only when it sets spec.hostname and spec.subdomain, and a headless Service with the same name as the subdomain exists in its namespace. StatefulSets do this for you through serviceName.
  • cluster.local is the default cluster domain. A task may use a different one, so check clusterDomain.

How short names work

A Pod in namespace sales gets roughly this /etc/resolv.conf:

nameserver 10.96.0.10
search sales.svc.cluster.local svc.cluster.local cluster.local
options ndots:5
  • pricing resolves via the first search domain (same namespace). pricing.ops resolves via svc.cluster.local (another namespace).
  • With ndots:5, any name with fewer than five dots is tried against the search list first, so external names like api.partner.net cost several failed lookups before the real one. A trailing dot (api.partner.net.) skips the search list.

Exam signal

"Reach Service X in namespace Y from a Pod in namespace Z" is answered with X.Y (or the full X.Y.svc.cluster.local). Plain X only works from inside namespace Y.

dnsPolicy and dnsConfig

dnsPolicyResolver the Pod uses
ClusterFirst (default)CoreDNS. Non-cluster names are forwarded upstream by CoreDNS
DefaultThe node's resolv.conf. Cluster names don't resolve. (Yes, Default isn't the default)
ClusterFirstWithHostNetCoreDNS for a hostNetwork: true Pod, which would otherwise fall back to Default
NoneOnly what you put in dnsConfig
spec:
  dnsPolicy: None
  dnsConfig:
    nameservers: ["10.96.0.10"]
    searches: ["sales.svc.cluster.local", "svc.cluster.local", "cluster.local"]
    options:
    - name: ndots
      value: "2"

For a few fixed names inside one Pod, spec.hostAliases adds entries to the Pod's /etc/hosts without touching DNS.

The Corefile

The kubeadm default, trimmed to the lines you'll touch (see k -n kube-system get cm coredns -o yaml for the rest):

.:53 {                                   # server block: every zone, port 53
    errors                               # log errors to stdout
    kubernetes cluster.local in-addr.arpa ip6.arpa {
       pods insecure                     # answer a-b-c-d.<ns>.pod names
       fallthrough in-addr.arpa ip6.arpa # pass unknown reverse lookups on
    }
    forward . /etc/resolv.conf           # non-cluster names go to the node's resolvers
    cache 30
    loop
    reload                               # re-read the Corefile when the ConfigMap changes
}

The full default also carries health, ready (probe endpoints), prometheus :9153 (metrics) and loadbalance (shuffles A records).

PluginDoes
kubernetesAnswers cluster names from the API. pods insecure enables the a-b-c-d.ns.pod records
forwardSends other queries upstream, here to the nameservers in the CoreDNS Pod's resolv.conf (the node's)
cacheCaches answers for up to 30 seconds
reloadPicks up Corefile changes without a restart (after the ConfigMap reaches the Pods)
loopDetects forwarding loops and stops CoreDNS (CrashLoopBackOff) rather than looping
hosts, rewriteStatic records, and rewriting one name to another
logLogs every query. Add it temporarily when debugging

Forwarding a private zone (stub domain)

Add a server block for the zone. Don't edit the forward . line, which would send all external lookups there.

corp.lan:53 {
    errors
    cache 30
    forward . 10.20.0.53
}
k -n kube-system edit configmap coredns          # add the block next to .:53
k -n kube-system rollout restart deploy coredns  # or wait for reload; restarting is faster and certain
k -n kube-system logs -l k8s-app=kube-dns --tail=20   # syntax errors show here

A typo in the Corefile

A broken Corefile can make CoreDNS crash on restart, and then every lookup in the cluster fails. After any edit, watch the Pods come back Ready and read their logs before moving on. Keep a copy: k -n kube-system get cm coredns -o yaml > coredns-backup.yaml.

Debugging a lookup

k run dns --rm -it --restart=Never --image=registry.k8s.io/e2e-test-images/jessie-dnsutils:1.3 -- \
  nslookup pricing.sales
k exec -n sales deploy/web -- cat /etc/resolv.conf      # right nameserver and search list?
k -n kube-system get pods -l k8s-app=kube-dns           # Running and Ready?
k -n kube-system get endpointslices -l kubernetes.io/service-name=kube-dns
k -n kube-system logs -l k8s-app=kube-dns
SymptomLikely cause
Every name fails, connection timed out; no servers could be reachedCoreDNS Pods down, kube-dns Service has no endpoints, or an egress NetworkPolicy blocks port 53
Cluster names fail, internet names workPod has dnsPolicy: Default, or wrong clusterDomain
Internet names fail, cluster names workUpstream in forward unreachable, or the node's resolv.conf is wrong
NXDOMAIN for a ServiceWrong namespace in the name, or the Service doesn't exist
kube-dns
Service name for CoreDNS, kept for compatibility.
ndots:5
Default option in Pod resolv.conf.
ClusterFirst
Default dnsPolicy.
svc.cluster.local
Suffix of every Service name in the default domain.

Scenarios

Scenario
A Pod in namespace web must call Service inventory in namespace stock. curl http://inventory fails with 'Could not resolve host', but CoreDNS is healthy. What is the simplest correct fix in the client config?
Scenario
Developers need names under build.internal to resolve from Pods using the company DNS server 10.30.4.4. All other external names must keep using the node's resolvers. What should you change?
Scenario
After someone edited the coredns ConfigMap, the CoreDNS Pods are in CrashLoopBackOff and their logs mention a detected loop. What is the most likely cause?

Drill

kubectl config use-context lab-dns

  1. Configure CoreDNS so names in corp.lan are resolved by 10.20.0.53, without changing how other names resolve.
  2. In namespace sales, find the IP that pricing.sales.svc.cluster.local resolves to and write it to /opt/answers/pricing-ip.txt.

Further reading

On this page