1. Introduction
A stuck Kubernetes Deployment rollout — where kubectl rollout status deployment/<name> hangs or returns Waiting for deployment rollout to finish: X out of N new replicas have been updated — blocks your CI/CD pipeline, leaves your cluster in a mixed state, and in some configurations can cause downtime if old pods are terminated before new ones become ready.
A rollout stalls when new pods can't become Ready. The old ReplicaSet waits to scale down, and the new ReplicaSet waits to scale up, because Kubernetes respects the maxUnavailable and maxSurge constraints. This means the root cause is almost always: the new pods aren't passing readiness probes.
2. What Rollout Stuck Means
Kubernetes rolls out a Deployment by creating a new ReplicaSet and gradually scaling it up while scaling the old one down. The pace is controlled by strategy.rollingUpdate.maxUnavailable (how many old pods can be removed at once) and maxSurge (how many extra pods can exist during rollout). If new pods don't become Ready, the rollout stops advancing because Kubernetes won't remove more old pods until new ones are serving traffic.
3. Common Causes
- New pods are in CrashLoopBackOff — application bug in the new version
- New pods have ImagePullBackOff — wrong image tag or registry auth failure
- New pods are stuck in ContainerCreating — volume mount or secret issue
- Readiness probe is failing — new version doesn't respond to the readiness check
- Resource limits are too tight — new pods are OOMKilled on startup
- ResourceQuota prevents new pods from being created
- progressDeadlineSeconds exceeded — rollout has been stuck longer than the configured deadline
- maxUnavailable: 0 and maxSurge: 0 (impossible configuration — rollout can never progress)
4. Step-by-Step Diagnosis and Fix
Step 1: Check rollout status and find the new pods
# Check the rollout
kubectl rollout status deployment/<name> -n <namespace> --timeout=30s
# If stuck, Ctrl+C and investigate
# Find the new ReplicaSet and its pods
kubectl get replicaset -n <namespace> -l <deployment-selector>
# Look for the RS with DESIRED > 0 but READY < DESIRED
# Get the new pods (highest generation)
kubectl get pods -n <namespace> -l <deployment-selector> --sort-by=.metadata.creationTimestamp
Step 2: Check what's wrong with the new pods
# Describe the new pods to see what's happening
kubectl describe pod <new-pod-name> -n <namespace>
# Check Events for: OOMKilled, ImagePullBackOff, ContainerCreating, Readiness failed
# Check pod status quickly
kubectl get pods -n <namespace> -l <selector>
# Look at STATUS and READY columns
# If CrashLoopBackOff: check container logs
kubectl logs <new-pod-name> -n <namespace> --previous
kubectl logs <new-pod-name> -n <namespace>
# If ImagePullBackOff: check image tag
kubectl get pod <new-pod-name> -n <namespace> -o jsonpath='{.spec.containers[0].image}'
Step 3: Rollback if the new version is broken
# Immediate rollback to previous version
kubectl rollout undo deployment/<name> -n <namespace>
# Watch the rollback complete
kubectl rollout status deployment/<name> -n <namespace>
# Rollback to a specific revision
kubectl rollout history deployment/<name> -n <namespace>
kubectl rollout undo deployment/<name> -n <namespace> --to-revision=<n>
Step 4: Fix and redeploy
# After fixing the root cause (image, config, resource limits):
# Update the deployment image:
kubectl set image deployment/<name> app=myregistry/myapp:fixed-tag -n <namespace>
# Force a rollout restart (uses same image, useful for config changes):
kubectl rollout restart deployment/<name> -n <namespace>
# Watch the new rollout:
kubectl rollout status deployment/<name> -n <namespace>
Step 5: Fix maxUnavailable / maxSurge constraints
# If the rollout is stuck because of overly conservative settings:
kubectl patch deployment <name> -n <namespace> --type='json' -p='[{"op":"replace","path":"/spec/strategy/rollingUpdate/maxUnavailable","value":"25%"},
{"op":"replace","path":"/spec/strategy/rollingUpdate/maxSurge","value":"25%"}]'
# Recommended settings for most deployments:
spec:
strategy:
type: RollingUpdate
rollingUpdate:
maxSurge: 1 # or 25%
maxUnavailable: 0 # zero-downtime: always add before removing
Step 6: Check ResourceQuota preventing new pods
# If the new pods can't be created at all:
kubectl get events -n <namespace> --field-selector reason=FailedCreate
# Look for: "exceeded quota"
# Check quota usage:
kubectl describe resourcequota -n <namespace>
# Fix: see /fix-kubernetes-resourcequota-exceeded.html
5. Verification Steps
# Successful rollout should complete without hanging
kubectl rollout status deployment/<name> -n <namespace>
# Expected: "deployment "my-app" successfully rolled out"
# All pods should be running the new image
kubectl get pods -n <namespace> -l <selector> -o jsonpath='{.items[*].spec.containers[0].image}'
# All should show the new tag
# No old ReplicaSet pods should remain
kubectl get replicaset -n <namespace> -l <selector>
# Old RS should show 0 DESIRED, 0 CURRENT
6. Common Mistakes
- Force-deleting new pods during a stuck rollout instead of reading their logs first
- Not using
--timeoutwithkubectl rollout statusin CI/CD — pipeline hangs indefinitely instead of failing - Setting
maxUnavailable: 0andmaxSurge: 0— this makes rollout impossible - Not checking
progressDeadlineSeconds— if it's set too low, healthy rollouts fail prematurely - Rolling back without investigating the root cause — you'll encounter the same issue next deployment
7. Prevention Tips
- Always use
kubectl rollout status --timeout=<n>sin CI/CD pipelines so stuck rollouts fail the pipeline rather than hanging - Set
progressDeadlineSecondsto a value that matches your expected deployment time (e.g. 300s for most apps) - Test readiness probe behavior in staging with the actual startup load before production rollouts
- Use image tags that include the git SHA or build ID — never use
:latestin production, it makes rollback impossible - Set up rollout notifications — alert when a rollout has been in progress for more than N minutes
- If your Deployment depends on a PVC, ensure PVC is bound before the rollout — a PVC in Pending will block the rollout
8. FAQ
The new pods are Running but the rollout is still stuck. Why?
Running doesn't mean Ready. The readiness probe may be failing — the pod is alive but not passing the health check that gates traffic routing. Check kubectl describe pod <new-pod> for "Readiness probe failed" events. The readiness probe may be configured too aggressively (timeout too short, wrong path) or the application genuinely isn't ready to serve traffic.
How do I cancel a stuck rollout without rolling back?
Pause the rollout: kubectl rollout pause deployment/<name>. This freezes the rollout in its current state — some pods running new version, some old. Then investigate and fix the issue, update the deployment spec, and resume: kubectl rollout resume deployment/<name>. Useful when you want to fix the new deployment config without reverting entirely.
9. Summary
| Stuck reason | Signal | Fix |
|---|---|---|
| New pods CrashLoopBackOff | pods STATUS = CrashLoopBackOff | Check logs; rollback; fix and redeploy |
| New pods ImagePullBackOff | pods STATUS = ImagePullBackOff | Fix image tag or registry credentials |
| Readiness probe failing | pods Running but 0/1 READY | Fix probe path/timeout or app startup behavior |
| ResourceQuota blocked | FailedCreate events | Free quota or increase limits |
| progressDeadline exceeded | Deployment condition: ProgressDeadlineExceeded | Rollback; increase deadline; fix root cause |
Explore More in This Category
Explore more in this category: Kubernetes guides. Browse all DevOps Compass articles or jump to: Kubernetes, AWS, CI/CD, Containers, Monitoring, Networking.