1. Introduction
A pod stuck in Pending is one of the more frustrating Kubernetes states to debug. If your pod is scheduled but keeps crashing instead, see Fix Kubernetes CrashLoopBackOff. Unlike CrashLoopBackOff or ImagePullBackOff, the container hasn't even started — the scheduler couldn't find a suitable node to place it on. The pod just sits there, waiting.
The Kubernetes scheduler evaluates every node against a set of constraints before it will place a pod. If no node passes all of those checks, the pod stays Pending indefinitely. The constraints range from available CPU and memory, to taints and tolerations, to node selectors and affinity rules, to PersistentVolumeClaim bindings.
This guide walks through every common reason a pod gets stuck in Pending with 'no nodes available' or similar scheduling failure messages, with the exact commands to identify the cause and the steps to fix it.
2. What 'Pod Pending' Actually Means
When you create a pod (directly or via a Deployment, StatefulSet, or Job), the kube-scheduler picks it up and tries to find a suitable node through two phases:
- Filtering — eliminates all nodes that fail any hard constraint (insufficient CPU/memory, taint mismatch, node selector mismatch, topology constraints)
- Scoring — ranks the remaining eligible nodes and picks the best one
If the filtering phase eliminates every node in the cluster, no scheduling decision can be made. The pod stays in Pending and the scheduler logs the reason in the pod's Events. That Events section is the primary diagnostic tool for this class of problem.
3. Common Causes
- Insufficient CPU or memory on all available nodes
- Resource requests set too high relative to node capacity
- All nodes are cordoned or the cluster has no Ready nodes
- Node taint with no matching toleration on the pod
- nodeSelector referencing a label that no node has
- nodeAffinity rule that no node satisfies
- PersistentVolumeClaim in Pending state or bound to a volume in the wrong zone
- Pod anti-affinity rules that conflict with existing pod placement
- Topology spread constraints that cannot be satisfied
- Namespace resource quota has been reached
- Cluster autoscaler is not enabled, or is enabled but the node group has hit its maximum
4. Step-by-Step Diagnosis and Fix
Step 1: Read the scheduler's reason from pod Events
This is always the first command to run. The Events section of kubectl describe pod tells you exactly why the scheduler rejected every node:
kubectl describe pod <pod-name> -n <namespace>
# Scroll to the Events section at the bottom:
# Events:
# Type Reason Age From Message
# ---- ------ ---- ---- -------
# Warning FailedScheduling 45s default-scheduler 0/3 nodes are available:
# 1 node(s) had untolerated taint {node-role.kubernetes.io/control-plane: }.
# 2 node(s) have insufficient memory.
# preemption: 0/3 nodes are eligible for preemption.
The message '0/3 nodes are available' followed by a breakdown per node is the scheduler's exact reasoning. Each line tells you how many nodes failed and why. Read all of them — a pod often fails multiple checks across different nodes.
Step 2: Check node capacity and resource pressure
Insufficient resources on all nodes is the most common cause of Pending pods in production, especially during traffic spikes or after a scale-down event:
# Check allocatable resources vs current requests on each node kubectl describe nodes | grep -A 8 'Allocated resources'
# Example output:
# Allocated resources:
# (Total limits may be over 100 percent, i.e., overcommitted.)
# Resource Requests Limits
# -------- -------- ------
# cpu 1850m (92%) 2200m (110%)
# memory 3.2Gi (89%) 4.1Gi (114%)
# Get a quick overview of all nodes kubectl get nodes kubectl top nodes # requires metrics-server
# See exactly how much is allocatable vs requested across the cluster kubectl get nodes -o custom-columns=\ 'NAME:.metadata.name,CPU:.status.allocatable.cpu,MEM:.status.allocatable.memory'
If nodes are at or near capacity, you have three options:
- Reduce the resource requests on the pod if they are over-specified for the actual workload
- Add nodes to the cluster (manually or by adjusting autoscaler node group limits)
- Evict or scale down other workloads to free up capacity
Step 3: Check node status and cordoning
A node that is cordoned or NotReady is excluded from scheduling entirely:
# Check node statuses kubectl get nodes
# Example showing a cordoned node:
# NAME STATUS ROLES AGE
# node-1 Ready worker 14d
# node-2 Ready,SchedulingDisabled worker 14d <-- cordoned
# node-3 NotReady worker 3m <-- not ready
# See why a node is NotReady kubectl describe node <node-name> | grep -A 10 'Conditions:'
# Uncordon a node that was manually cordoned kubectl uncordon <node-name>
A NotReady node typically means the kubelet has stopped reporting to the API server — the node may be down, network-partitioned, or under severe resource pressure. Inspect the node directly if possible, and check cloud provider health checks if running on AWS, GCP, or Azure.
Step 4: Check for taint and toleration mismatches
Taints mark nodes as unsuitable for pods that do not explicitly tolerate them. This is commonly used for dedicated node groups (GPU nodes, spot instances, system components):
# List all taints on all nodes kubectl get nodes -o custom-columns=\ 'NAME:.metadata.name,TAINTS:.spec.taints'
# Example output showing a taint:
# NAME TAINTS
# node-1 [map[effect:NoSchedule key:dedicated value:gpu]]
# node-2 [map[effect:NoSchedule key:node-role.kubernetes.io/control-plane]]
# Check what tolerations the pending pod has kubectl get pod <pod-name> -n <namespace> -o jsonpath='{.spec.tolerations}'
If your pod needs to run on a tainted node, add the matching toleration to the pod spec:
# Pod spec tolerations example spec: tolerations: - key: "dedicated" operator: "Equal" value: "gpu" effect: "NoSchedule"
# To tolerate a key regardless of value: tolerations: - key: "dedicated" operator: "Exists" effect: "NoSchedule"
Step 5: Check nodeSelector and node labels
A nodeSelector in the pod spec restricts scheduling to nodes with specific labels. If no node has that label, the pod can never be scheduled:
# Check the pod's nodeSelector kubectl get pod <pod-name> -n <namespace> -o jsonpath='{.spec.nodeSelector}'
# Example output:
# {"disktype":"ssd","zone":"us-east-1a"}
# List nodes and their labels to find a match kubectl get nodes --show-labels
# Filter for nodes with a specific label kubectl get nodes -l disktype=ssd
# Add a label to a node (if appropriate): kubectl label node <node-name> disktype=ssd
If the nodeSelector references a label that was accidentally removed or never applied to the new node group, either add the label to the correct nodes or update the pod spec to remove the overly restrictive selector.
Step 6: Check nodeAffinity rules
nodeAffinity is a more expressive version of nodeSelector. Required affinity rules (requiredDuringSchedulingIgnoredDuringExecution) act as hard constraints — if they can't be satisfied, the pod won't schedule:
# Check affinity rules on the pending pod kubectl get pod <pod-name> -n <namespace> \ -o jsonpath='{.spec.affinity}' | jq .
# Example required affinity that may be too restrictive:
# affinity:
# nodeAffinity:
# requiredDuringSchedulingIgnoredDuringExecution:
# nodeSelectorTerms:
# - matchExpressions:
# - key: topology.kubernetes.io/zone
# operator: In
# values:
# - us-east-1c # no nodes exist in this zone
# Check what zone labels your nodes have kubectl get nodes -L topology.kubernetes.io/zone
For zone-specific affinity rules, verify that nodes exist in the referenced zone. This is a common failure point after a node group in one availability zone is scaled to zero or terminated by the cloud provider.
Step 7: Check PersistentVolumeClaim status
A pod with a volumeClaimTemplate or explicit PVC reference will stay Pending if the PVC is not yet bound — the scheduler won't place the pod until the volume is ready:
# Check PVC status in the namespace kubectl get pvc -n <namespace>
# Example showing an unbound PVC:
# NAME STATUS VOLUME CAPACITY ACCESS MODES STORAGECLASS
# data-pvc Pending <none> <none> <none> gp2
# Describe the PVC to see why it's Pending kubectl describe pvc <pvc-name> -n <namespace>
# Common messages:
# 'no persistent volumes available for this claim'
# 'waiting for first consumer to be created before binding' (WaitForFirstConsumer)
Common PVC Pending causes and fixes:
- StorageClass does not exist — verify with: kubectl get storageclass
- No PersistentVolumes available for static provisioning — create a PV or switch to dynamic provisioning
- Volume binding mode is WaitForFirstConsumer — this is intentional; the PVC binds once the pod is scheduled, so the pod and PVC can appear mutually blocked. The scheduler resolves this automatically; it just needs a node that satisfies all other constraints first
- Cross-zone volume binding — the PVC is bound to a volume in a different availability zone than the nodes where the pod can be scheduled. Fix by ensuring node affinity rules and storage topology are aligned
Step 8: Check namespace resource quotas
If a ResourceQuota is defined on the namespace and the new pod would exceed it, the pod is rejected before it even reaches the scheduler:
# Check resource quotas in the namespace kubectl describe resourcequota -n <namespace>
# Example output:
# Name: compute-quota
# Resource Used Hard
# -------- ---- ----
# requests.cpu 1900m 2000m <-- nearly full
# requests.memory 3.8Gi 4Gi <-- nearly full
# pods 9 10 <-- at limit
# The pod Events will also show:
# 'exceeded quota: compute-quota, requested: pods=1,
# used: pods=10, limited: pods=10'
To resolve a quota issue:
- Increase the quota limit if capacity exists on the nodes (kubectl edit resourcequota -n <namespace>)
- Delete unused pods or reduce replica counts in the namespace to free up quota
- Review whether the quota limits are still appropriate for the workload — quotas set during initial cluster setup are often not revisited as workloads grow
Step 9: Check cluster autoscaler (if enabled)
If your cluster uses the Kubernetes Cluster Autoscaler, it should automatically add nodes when pods are unschedulable. If pods remain Pending for more than a few minutes, the autoscaler may be blocked:
# Check cluster autoscaler logs kubectl logs -n kube-system deployment/cluster-autoscaler | tail -50
# Common autoscaler messages to look for:
# 'pod is unschedulable' — autoscaler has noticed the pending pod
# 'node group has reached maximum size' — scale-out is blocked
# 'no expansion possible' — no node group can satisfy the pod's constraints
# 'scale up in progress' — autoscaler is adding nodes, wait
# Check current node group sizes (AWS EKS example) aws autoscaling describe-auto-scaling-groups \ --query 'AutoScalingGroups[*].{Name:AutoScalingGroupName, Min:MinSize,Max:MaxSize,Desired:DesiredCapacity}'
If the autoscaler shows 'node group has reached maximum size', you need to increase the maximum node count for the relevant node group in your cloud provider's autoscaling configuration.
5. Verification Steps
Once you've identified and addressed the root cause, verify the pod moves out of Pending:
# Watch the pod status update in real time kubectl get pod <pod-name> -n <namespace> -w
# Expected recovery:
# NAME READY STATUS RESTARTS AGE
# myapp-abc 0/1 Pending 0 8m
# myapp-abc 0/1 Pending 0 8m <-- still evaluating
# myapp-abc 0/1 Init:0/1 0 8m
# myapp-abc 0/1 Running 0 8m
# myapp-abc 1/1 Running 0 8m <-- scheduled and healthy
# Check which node the pod was scheduled onto kubectl get pod <pod-name> -n <namespace> -o wide
# Confirm no scheduling events remain kubectl describe pod <pod-name> -n <namespace> | grep -A 20 'Events:'
6. Common Mistakes
- Only looking at one line of the scheduling failure message — the full breakdown shows all the reasons across all nodes, and you need to address them all
- Adding nodes to the cluster without fixing a nodeSelector or taint issue — new nodes won't help if the pod can't be placed on them either
- Right-sizing memory limits but not requests — Kubernetes schedules on requests, not limits
- Confusing WaitForFirstConsumer PVC behaviour with a real Pending problem — wait a minute and check if it self-resolves
- Removing a node taint to fix one pod, inadvertently opening a dedicated node group to general workloads
- Not checking resource quotas — a quota rejection appears in Events but can be easy to miss if you stop reading after seeing resource pressure messages
- Assuming the cluster autoscaler will fix everything — it can't add nodes that satisfy impossible constraints
7. Prevention Tips
- Set resource requests accurately — use Vertical Pod Autoscaler (VPA) in recommendation mode to get data-driven request sizing
- Use Horizontal Pod Autoscaler (HPA) to scale workloads before nodes saturate, rather than after
- Enable the cluster autoscaler with sensible max limits and verify it can scale the right node groups for your workload profiles
- Audit nodeSelectors and node affinity rules when making infrastructure changes — removing a label from a node group silently breaks pods that depend on it
- Monitor PodSchedulingLatency in Prometheus or use kube-state-metrics to alert on pods in Pending state longer than 5 minutes. Once scheduled, pods may fail with ImagePullBackOff if registry credentials or image references are wrong.
- Review namespace ResourceQuotas periodically as workloads grow — quotas set at cluster creation are rarely revisited
- Document which node groups carry which taints and labels — this is critical operational knowledge that frequently lives only in one engineer's head
- Use kubectl explain pod.spec.affinity and dry-run deployments to validate scheduling rules before applying them to production
8. Summary
A pod stuck in Pending means the scheduler couldn't find a node that satisfies all of the pod's requirements. Start with kubectl describe pod and read the full Events section — the scheduler tells you exactly which nodes failed and why.
| Scheduling failure reason | Where to look | Fix |
|---|---|---|
| Insufficient CPU / memory | kubectl describe nodes (Allocated resources) | Reduce requests, add nodes, or evict low-priority pods |
| Node cordoned / NotReady | kubectl get nodes | Uncordon node or replace unhealthy node |
| Taint not tolerated | kubectl get nodes -o custom-columns (TAINTS) | Add toleration to pod spec |
| nodeSelector no match | kubectl get nodes --show-labels | Add label to node or relax nodeSelector |
| nodeAffinity no match | kubectl get pod -o jsonpath (affinity) | Update affinity values or add nodes in required zone |
| PVC Pending | kubectl get pvc -n <namespace> | Fix StorageClass, PV, or cross-zone topology issue |
| ResourceQuota exceeded | kubectl describe resourcequota | Increase quota or free up namespace capacity |
| Autoscaler at max | cluster-autoscaler logs | Increase node group maximum in cloud provider config |
Most Pending issues come down to resource pressure or a constraint mismatch. Documenting your resolution steps in a DevOps runbook helps the next on-call engineer resolve the same issue faster. The scheduler's failure message in kubectl describe pod is specific enough to point you directly at the cause — don't skip reading it in full.