1. Introduction
DNS resolution failures in Kubernetes can cause a wide range of symptoms that look completely unrelated to DNS: services that can't find each other, applications that fail to connect to databases, HTTP clients that hang indefinitely, and pods that appear healthy but can't reach anything by hostname. When DNS breaks in a cluster, the knock-on effects are significant because nearly all inter-service communication relies on it.
Kubernetes uses CoreDNS (since 1.13) to handle DNS for pods. If CoreDNS is degraded, if pods have incorrect resolv.conf, or if the cluster's DNS policies are misconfigured, service discovery fails. This guide covers the full diagnostic path from confirming DNS is broken through to fixing CoreDNS, ndots configuration, and search domain issues.
2. How Kubernetes DNS Works
Every pod in a Kubernetes cluster gets a /etc/resolv.conf injected by the kubelet. By default it looks like:
nameserver 10.96.0.10 # CoreDNS ClusterIP (kube-dns Service)
search default.svc.cluster.local svc.cluster.local cluster.local
options ndots:5
This means pod DNS queries go to CoreDNS first. CoreDNS resolves cluster-internal names (like my-service.my-namespace.svc.cluster.local) and forwards external names to the upstream DNS configured in the cluster (typically the VPC DNS or a public resolver).
3. Common Causes
- CoreDNS pods are not running or are OOMKilled
- CoreDNS ConfigMap has incorrect upstream forwarders
- Pod's
dnsPolicyis set toNoneorDefaultwhen it should beClusterFirst - Network policy blocking UDP/TCP port 53 between pods and CoreDNS
- Custom
dnsConfigon the pod that overrides the cluster DNS - The
kube-dnsService doesn't have endpoints pointing to CoreDNS pods - CoreDNS is running but upstream DNS resolution is failing (VPC DNS or internet DNS unreachable)
- CoreDNS ConfigMap's
forwardplugin pointing to an unreachable nameserver - High DNS query rate causing CoreDNS to be overwhelmed (common in large clusters)
4. Step-by-Step Diagnosis and Fix
Step 1: Confirm DNS is actually broken
# Run a DNS test pod in the affected namespace
kubectl run dns-test --image=busybox:1.35 --rm -it --restart=Never -- sh
# Inside the pod:
nslookup kubernetes.default # should return the API server IP
nslookup kubernetes.default.svc.cluster.local # FQDN form
nslookup google.com # external DNS
# Test with a specific nameserver:
nslookup kubernetes.default 10.96.0.10 # query CoreDNS directly
# Check /etc/resolv.conf inside the pod:
cat /etc/resolv.conf
Step 2: Check CoreDNS pod health
# Check CoreDNS pod status
kubectl get pods -n kube-system -l k8s-app=kube-dns
# If pods are CrashLoopBackOff or not Running:
kubectl describe pod -n kube-system -l k8s-app=kube-dns
kubectl logs -n kube-system -l k8s-app=kube-dns --tail=50
# Check if CoreDNS is being OOMKilled (common cause of DNS flaps)
kubectl get pod -n kube-system -l k8s-app=kube-dns -o jsonpath='{.items[*].status.containerStatuses[*].lastState}'
# If OOMKilled, increase CoreDNS memory limit:
kubectl edit deployment coredns -n kube-system
# Increase resources.limits.memory from 170Mi to 300Mi+
Step 3: Verify the kube-dns Service and endpoints
# Check the kube-dns Service
kubectl get service kube-dns -n kube-system
# NAME TYPE CLUSTER-IP EXTERNAL-IP PORT(S)
# kube-dns ClusterIP 10.96.0.10 <none> 53/UDP,53/TCP
# Check it has endpoints pointing to CoreDNS pods
kubectl get endpoints kube-dns -n kube-system
# ENDPOINTS should NOT be <none>
# If endpoints are empty:
kubectl describe service kube-dns -n kube-system
# Look at Selector — should match CoreDNS pod labels
Step 4: Check CoreDNS ConfigMap for upstream issues
# View the CoreDNS configuration
kubectl get configmap coredns -n kube-system -o yaml
# A typical healthy Corefile looks like:
# .:53 {
# errors
# health { lameduck 5s }
# ready
# kubernetes cluster.local in-addr.arpa ip6.arpa {
# pods insecure
# fallthrough in-addr.arpa ip6.arpa
# }
# prometheus :9153
# forward . /etc/resolv.conf { <-- forwards external queries to VPC DNS
# max_concurrent 1000
# }
# cache 30
# loop
# reload
# loadbalance
# }
# If the forward directive points to unreachable DNS servers:
# Change to use a reliable upstream:
forward . 8.8.8.8 8.8.4.4 {
max_concurrent 1000
}
Step 5: Check for network policies blocking DNS
"# List network policies in the affected namespace
kubectl get networkpolicy -n <namespace>
# A restrictive NetworkPolicy can block UDP 53 traffic from pods to CoreDNS
# The following policy allows DNS (add to your namespace's policy):
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: allow-dns
namespace: <your-namespace>
spec:
podSelector: {}
egress:
- ports:
- protocol: UDP
port: 53
- protocol: TCP
port: 53
policyTypes:
- Egress
Step 6: Fix pod DNS policy misconfiguration
# Check what dnsPolicy your pods are using
kubectl get pod <pod-name> -n <namespace> -o jsonpath='{.spec.dnsPolicy}'
# ClusterFirst (default) = use CoreDNS
# Default = use node's /etc/resolv.conf (not cluster DNS)
# None = use only dnsConfig
# If set to Default or None inadvertently, fix in the Deployment spec:
spec:
template:
spec:
dnsPolicy: ClusterFirst # restore to default
Step 7: Fix ndots and search domain issues for external DNS
# If external domain lookups are slow or failing due to ndots:5,
# override per-deployment with dnsConfig:
spec:
template:
spec:
dnsPolicy: ClusterFirst
dnsConfig:
options:
- name: ndots
value: "1" # reduces unnecessary search domain iterations
- name: timeout
value: "2"
- name: attempts
value: "3"
5. Verification Steps
# After fixing, run the DNS test again:
kubectl run dns-verify --image=busybox:1.35 --rm -it --restart=Never -- sh
# Inside:
nslookup kubernetes.default # should resolve
nslookup my-service.my-namespace.svc.cluster.local # your service
nslookup google.com # external
# Check CoreDNS metrics (if Prometheus is available):
kubectl port-forward -n kube-system service/kube-dns 9153:9153
curl http://localhost:9153/metrics | grep coredns_dns_request_duration
6. Common Mistakes
- Not checking CoreDNS memory usage — OOMKilled CoreDNS pods are a common cause of intermittent DNS failures
- Forgetting to check NetworkPolicy egress rules — a policy that restricts all egress silently breaks DNS
- Changing the CoreDNS ConfigMap without testing with
kubectl rollout restart— changes aren't applied until CoreDNS reloads - Using
dnsPolicy: Defaultin pods when you need cluster-internal service discovery - Not considering ndots when debugging slow external queries — high ndots causes multiple unnecessary DNS lookups before the real query
7. Prevention Tips
- Monitor CoreDNS memory and CPU usage — alert before it gets OOMKilled in production
- Set appropriate resource limits on CoreDNS that match your cluster's query volume
- For large clusters, deploy more CoreDNS replicas and distribute them across nodes
- If pods can't reach external services, also check EKS pod internet access — DNS may be working but NAT Gateway or route tables may be missing
- Always test DNS in new namespaces as part of namespace bootstrap validation
- Use Service discovery debugging steps alongside DNS checks — communication failures combine both DNS and network connectivity
8. FAQ
My pods can reach Services by IP but not by name. Is this a DNS issue?
Yes — classic symptom. The network connectivity is fine, but DNS resolution is broken. Start with nslookup kubernetes.default from inside the pod. If that fails, CoreDNS isn't reachable. If it succeeds for kubernetes.default but fails for your service name, there might be a namespace issue — try the fully qualified name: nslookup my-service.my-namespace.svc.cluster.local.
DNS works for some pods but not others in the same cluster. Why?
Check: (1) Is there a NetworkPolicy in the affected namespace that blocks egress to port 53? (2) Do the affected pods have a different dnsPolicy? (3) Are the affected pods on a specific node where the node's iptables rules are corrupted? Run the DNS test pod in the exact same namespace and node as the failing pods.
External DNS queries are very slow (2-3 seconds) even though CoreDNS is healthy. What's happening?
This is almost always the ndots:5 problem. With 5 search domains and ndots:5, a query for google.com generates 6 DNS queries before getting a result (one for each search domain, then the bare query). Reduce ndots to 1 or 2 with dnsConfig.options in your pod spec. The search domains are still used for short names like my-service.
9. Summary
| Symptom | Cause | Fix |
|---|---|---|
| All DNS fails in pod | CoreDNS down or unreachable | Check CoreDNS pods; check kube-dns endpoints |
| Cluster DNS fails, external works | CoreDNS misconfigured | Review CoreDNS ConfigMap kubernetes plugin |
| Intermittent DNS failures | CoreDNS OOMKilled | Increase CoreDNS memory limits |
| DNS fails in one namespace only | NetworkPolicy blocking port 53 | Add egress allow rule for UDP/TCP 53 |
| External DNS slow (2-3s) | ndots:5 causing extra lookups | Set ndots:1 in pod dnsConfig |
Explore More in This Category
Explore more in this category: Kubernetes guides. Browse all DevOps Compass articles or jump to: Kubernetes, AWS, CI/CD, Containers, Monitoring, Networking.