1. Introduction

A NotReady node is a cluster-level emergency. Any pods currently running on that node may be unreachable or terminated. New pods won't be scheduled to it. If enough nodes go NotReady simultaneously, your cluster loses capacity and workloads fail entirely. The cause can be a failed kubelet, a broken network plugin, resource exhaustion, or loss of connectivity to the control plane.

This guide walks through the complete diagnostic path: from what kubectl get nodes tells you, through SSH-level investigation, to resolution and pod rescheduling.

2. What NotReady Actually Means

A node is Ready when the kubelet reports all of the following: sufficient disk, memory, and PID resources; the container runtime is responsive; the network plugin is configured; and the node can communicate with the control plane. If any of these conditions fail for more than node-monitor-grace-period (default 40s), the node transitions to NotReady.

3. Common Causes

4. Step-by-Step Diagnosis and Fix

Step 1: Read the node conditions

# Check all node conditions
kubectl describe node <node-name> | grep -A 20 "Conditions:"

# Look for:
# MemoryPressure:   False  (True = low memory)
# DiskPressure:     False  (True = low disk)
# PIDPressure:      False  (True = too many processes)
# Ready:            False  (this is the main signal)
# NetworkUnavailable: False  (True = CNI failure)

# Also check Events at the bottom:
kubectl describe node <node-name> | tail -20

Step 2: Check the kubelet on the node

# SSH into the node (or use SSM Session Manager for AWS)
ssh -i key.pem ec2-user@<node-ip>

# Check kubelet service status
systemctl status kubelet
journalctl -u kubelet -n 100 --no-pager

# Common kubelet errors:
# "failed to get node info" = API server unreachable
# "certificate has expired" = TLS cert issue
# "runtime not responding" = containerd/Docker down
# "node disk pressure" = disk full

# Restart kubelet if it's stopped
systemctl restart kubelet
systemctl enable kubelet

Step 3: Check container runtime

# Check containerd
systemctl status containerd
journalctl -u containerd -n 50 --no-pager

# Test if containerd responds
crictl info

# Check Docker (if used)
systemctl status docker
docker info

# Restart if needed
systemctl restart containerd
# Or: systemctl restart docker

Step 4: Check CNI plugin

# From the control plane: check CNI DaemonSet pods on the affected node
NODE=<node-name>
kubectl get pods -n kube-system   --field-selector spec.nodeName=$NODE | grep -E "calico|flannel|aws-node|cilium|weave"

# Check CNI pod logs
kubectl logs -n kube-system <cni-pod> --tail=50

# On the node itself: check CNI config files
ls /etc/cni/net.d/
cat /etc/cni/net.d/*.conf

# Restart CNI pod to force re-initialization:
kubectl delete pod -n kube-system <cni-pod-on-node>

Step 5: Check control plane connectivity from the node

# On the node, test connectivity to the API server
# Get API server endpoint:
kubectl cluster-info 2>/dev/null | grep "Kubernetes master\|Kubernetes control plane"

# Test from the node:
curl -k https://<api-server-endpoint>:6443/healthz
# Expected: ok

# Check for network issues to the API server:
traceroute <api-server-ip>

# For EKS: verify the node can reach the VPC endpoint
nslookup <eks-cluster-endpoint>

Step 6: Check disk and memory

# On the node:
df -h           # disk usage — look for 100% on any filesystem
free -m         # memory
df -i           # inode usage

# Find what's using disk:
du -sh /var/lib/docker/* 2>/dev/null | sort -rh | head -10
du -sh /var/log/containers/* | sort -rh | head -10

# Emergency disk cleanup:
docker system prune -af 2>/dev/null || crictl rmi --prune 2>/dev/null
journalctl --vacuum-size=200M

Step 7: Drain, fix, and return node to service

# Cordon to prevent new pods scheduling during repair
kubectl cordon <node-name>

# Drain existing pods (if node will be rebooted)
kubectl drain <node-name> --ignore-daemonsets --delete-emptydir-data --force

# After fixing: uncordon to allow scheduling again
kubectl uncordon <node-name>

# Watch node recover
kubectl get nodes -w

5. Verification Steps

# Node should show Ready
kubectl get nodes
# NAME              STATUS   ROLES    AGE   VERSION
# node-1            Ready    <none>   10d   v1.29.x

# All node conditions should be False except Ready
kubectl describe node <node-name> | grep -A 8 "Conditions:"

# Pods should reschedule and become Running
kubectl get pods -A --field-selector spec.nodeName=<node-name>

6. Common Mistakes

7. Prevention Tips

8. FAQ

Multiple nodes went NotReady at the same time. What happened?

Simultaneous NotReady across multiple nodes almost always indicates: (1) a control plane failure (API server, etcd), (2) a network partition between the nodes and control plane, or (3) a cluster-wide issue like a certificate expiry. Check the control plane components first: kubectl get pods -n kube-system. If the control plane is unreachable, SSH directly to a node and check journalctl -u kubelet.

The node shows Ready again but pods aren't rescheduling. Why?

Pods from a recovered NotReady node are rescheduled by the controller manager, but only if the pods have a managing controller (Deployment, ReplicaSet, etc.). Standalone pods are not automatically rescheduled. Also check that the node isn't cordoned (kubectl get nodes shows SchedulingDisabled) — run kubectl uncordon <node> to allow scheduling.

9. Summary

Condition showingRoot causeFix
NetworkUnavailable=TrueCNI plugin failureRestart CNI DaemonSet pod on node
MemoryPressure=TrueNode out of memoryEvict pods; add capacity; right-size workloads
DiskPressure=TrueNode disk fullClean logs/images; enable log rotation
kubelet stoppedProcess crashsystemctl restart kubelet on the node
API server unreachableNetwork or control plane issueCheck VPC routes, SGs, control plane health

Explore More in This Category

Explore more in this category: Kubernetes guides. Browse all DevOps Compass articles or jump to: Kubernetes, AWS, CI/CD, Containers, Monitoring, Networking.