1. Introduction
A readiness probe failure in Kubernetes doesn't kill your container. It does something subtler: it removes the pod from the endpoints of every Service that targets it, so traffic stops being routed to that pod until the probe passes again. From a user perspective this can cause intermittent request failures, a sudden reduction in serving capacity, or a complete loss of traffic to a deployment — depending on how many pods fail at once.
Because readiness failures don't cause visible restarts, they're easier to miss than CrashLoopBackOff or liveness failures. This guide covers how to diagnose a readiness probe failure from first principles, fix the most common causes, and design probes that behave correctly under real production conditions.
2. How Readiness Probes Work
When a readiness probe fails, Kubernetes sets the pod's Ready condition to False and removes it from the Endpoints object for all Services that select it. Traffic stops reaching the pod. When the probe starts passing again, the pod is added back to the endpoints and traffic resumes.
This is intentionally different from what a liveness failure does:
| Probe type | On failure | Pod restarted? | Traffic removed? |
|---|---|---|---|
| Liveness | Container killed and restarted | Yes | During restart only |
| Readiness | Pod removed from Service endpoints | No | Yes, until probe passes |
| Startup | Container killed if not passing by deadline | Yes | Yes, during startup |
Readiness probes use the same three mechanisms as liveness probes (HTTP GET, TCP Socket, Exec) and the same timing fields (initialDelaySeconds, periodSeconds, timeoutSeconds, failureThreshold, successThreshold). The key difference is that successThreshold can be greater than 1 for readiness probes — useful for requiring a pod to pass the check multiple times in a row before receiving traffic.
3. Common Causes
- Application hasn't finished initialising yet — the readiness endpoint returns 503 during warmup
initialDelaySecondstoo low — the probe fires before the app is ready to serve traffic- A downstream dependency (database, cache, message queue) is unavailable and the readiness check reports this correctly
- Application is under high load and the readiness endpoint is too slow to respond within
timeoutSeconds - Wrong path or port in the probe spec — consistently returns 404 or connection refused
- The readiness endpoint is the same as the liveness endpoint and includes checks that should only affect traffic routing, not container lifetime
- A configuration reload or graceful shutdown puts the app in a temporarily not-ready state
- Sidecar container (e.g. Envoy/Istio proxy) not yet ready when the main container's readiness probe fires
4. Step-by-Step Diagnosis and Fix
Step 1: Identify that a readiness probe failure is the problem
Readiness failures show up differently from other pod issues. The pod may appear to be Running but receive no traffic:
# Check pod Ready status — 0/1 means not ready, 1/1 means ready
kubectl get pod <pod-name> -n <namespace>
# Example showing a not-ready pod:
# NAME READY STATUS RESTARTS AGE
# myapp-7b9f4 0/1 Running 0 4m
# Note: STATUS=Running but READY=0/1 — the container is alive but
# the readiness probe is failing. This is the readiness probe pattern.
# CrashLoopBackOff or Error in STATUS = different problem entirely.
# Confirm with kubectl describe and look for probe failure events:
kubectl describe pod <pod-name> -n <namespace>
# Events section will show:
# Warning Unhealthy 30s kubelet Readiness probe failed:
# HTTP probe failed with statuscode: 503
# Check whether the pod is in the Service endpoints:
kubectl get endpoints <service-name> -n <namespace>
# A healthy pod will have its IP listed here.
# A pod failing readiness will be absent or show NotReadyAddresses.
Step 2: Check the readiness probe configuration
# Get the current readiness probe spec
kubectl get pod <pod-name> -n <namespace> -o jsonpath='{.spec.containers[0].readinessProbe}' | jq .
# Example output:
# {
# "httpGet": {
# "path": "/ready",
# "port": 8080
# },
# "initialDelaySeconds": 5,
# "periodSeconds": 10,
# "timeoutSeconds": 1,
# "failureThreshold": 3,
# "successThreshold": 1
# }
Step 3: Test the readiness endpoint manually
Always test the actual probe target before changing configuration. This tells you whether the endpoint logic is wrong or the timing is wrong:
# Test the readiness endpoint from inside the container
kubectl exec -it <pod-name> -n <namespace> -- curl -sv http://localhost:8080/ready
# Note the response: status code, body, response time.
# If it returns 503 with a body like {"status":"DOWN","db":"UNAVAILABLE"},
# the problem is a downstream dependency, not probe configuration.
# If it returns 404, the path is wrong.
# If the request hangs, the port is wrong or the app is blocked.
# Check which port the app is listening on:
kubectl exec -it <pod-name> -n <namespace> -- ss -tlnp
# Time the response to check against timeoutSeconds:
kubectl exec -it <pod-name> -n <namespace> -- time curl -s http://localhost:8080/ready > /dev/null
Step 4: Fix — Application not ready during startup
If the readiness probe fails immediately after pod start but the application eventually becomes healthy, initialDelaySeconds is too low for your startup time:
# Check how long the app takes to reach a ready state.
# In the Events, note the timestamp of the first Unhealthy event
# and compare it to when the container started.
# Fix: increase initialDelaySeconds to cover startup time
readinessProbe:
httpGet:
path: /ready
port: 8080
initialDelaySeconds: 30 # Enough time for the app to initialise
periodSeconds: 10
timeoutSeconds: 3
failureThreshold: 3
successThreshold: 1
# Better fix for variable startup: use a startupProbe to hold off
# both liveness AND readiness until startup completes:
startupProbe:
httpGet:
path: /ready
port: 8080
failureThreshold: 30 # Up to 5 minutes to start
periodSeconds: 10
Step 5: Fix — Downstream dependency causing readiness failure
This is the most nuanced readiness probe issue. If the /ready endpoint returns 503 because a database or external service is unreachable, the behaviour is actually correct — the pod should stop receiving traffic if it can't successfully handle requests. But you need to understand whether this is intentional:
# Check whether the dependency failure is expected or unexpected:
kubectl exec -it <pod-name> -n <namespace> -- curl -sv http://localhost:8080/ready
# Response body example (Spring Boot Actuator):
# {
# "status": "DOWN",
# "components": {
# "db": {
# "status": "DOWN",
# "details": { "error": "connection refused to postgres:5432" }
# }
# }
# }
# If the database is legitimately down — fix the database, not the probe.
# If the database is fine but the connection string is wrong:
kubectl get secret <db-secret> -n <namespace> -o jsonpath='{.data}' | jq 'to_entries | map({key, value: (.value | @base64d)})'
# If you want the app to receive traffic even when the DB is down
# (e.g. to serve cached or static content), remove the DB check
# from the readiness endpoint and add it to a separate monitoring check.
Step 6: Fix — Probe timeout too low under load
If the readiness endpoint responds correctly in normal operation but times out when the application is under load, increase timeoutSeconds and consider making the readiness endpoint less dependent on external calls:
readinessProbe:
httpGet:
path: /ready
port: 8080
initialDelaySeconds: 20
periodSeconds: 10
timeoutSeconds: 5 # Allow 5 seconds for the endpoint to respond
failureThreshold: 3 # 3 consecutive failures before removing from endpoints
successThreshold: 1 # 1 success to add back to endpoints
# For successThreshold > 1 (require multiple passes before accepting traffic):
# Useful when a pod coming out of a bad state should be stable for a
# few probe cycles before receiving traffic again.
successThreshold: 2 # Must pass twice in a row before getting traffic
Step 7: Fix — Istio / service mesh sidecar not ready
If you're running Istio or another service mesh, the Envoy sidecar proxy may not be ready when your application's readiness probe fires, causing failures even when your application is healthy:
# Check if Istio injection is enabled and the sidecar is present:
kubectl describe pod <pod-name> -n <namespace> | grep -A 5 "Containers:"
# Look for istio-proxy or envoy alongside your main container
# Istio 1.7+ handles this with holdApplicationUntilProxyStarts:
# In your IstioOperator or mesh config:
# meshConfig:
# defaultConfig:
# holdApplicationUntilProxyStarts: true
# Or add an initContainer to wait for the proxy to be ready:
initContainers:
- name: wait-for-proxy
image: busybox
command: ['sh', '-c', 'until wget -qO- http://localhost:15021/healthz/ready; do sleep 2; done']
5. Verification Steps
After applying your fix, confirm the pod becomes Ready and stays Ready:
# Watch the Ready status column update
kubectl get pod <pod-name> -n <namespace> -w
# Expected: READY transitions from 0/1 to 1/1 and stays there
# Confirm the pod is in the Service endpoints
kubectl get endpoints <service-name> -n <namespace>
# Expected: pod IP appears in the Addresses column (not NotReadyAddresses)
# Verify no new readiness failure events
kubectl get events -n <namespace> --field-selector reason=Unhealthy --sort-by='.lastTimestamp' | tail -10
# For a Deployment, verify the rollout considers all pods ready:
kubectl rollout status deployment/<name> -n <namespace>
# Expected: deployment "my-app" successfully rolled out
6. Common Mistakes
- Using the same endpoint for liveness and readiness — they serve different purposes. Liveness = process alive. Readiness = ready to serve traffic now. A pod can be alive but legitimately not ready (e.g. warming up a cache).
- Making the readiness probe check every downstream dependency — this turns any dependency outage into a traffic-routing crisis even when the app is healthy
- Not realising the pod is failing readiness because the STATUS column still shows
Running— always check the READY column, not just STATUS - Setting
successThresholdtoo high — if a pod needs to pass the probe 5 times before receiving traffic, rolling deployments will be extremely slow - Not accounting for Istio sidecar startup when the readiness probe fires in the first few seconds
- Assuming a readiness failure is the same as a liveness failure — the pod is not restarted, it just stops receiving traffic. Check the Events section carefully to distinguish the two.
7. Prevention Tips
- Design your
/readyendpoint to represent "can I handle a request right now" — not "is everything I depend on healthy" - Use
startupProbefor slow-starting applications soinitialDelaySecondsdoesn't need to be set to a worst-case value - Set
progressDeadlineSecondson Deployments (e.g. 300s) so a stalled rollout caused by a failing readiness probe is detected and surfaces as a deployment failure rather than hanging indefinitely - Monitor
kube_endpoint_address_not_readyin Prometheus — this metric increments every time a pod is removed from Service endpoints due to readiness failure - Test readiness probes under load in staging — an endpoint that responds in 200ms at idle may timeout under production traffic volume
- Use separate paths for liveness and readiness:
/healthz(liveness — just process alive) and/ready(readiness — ready to serve) - If using liveness probes, set
failureThresholdhigher on readiness than liveness — it's safer to briefly stop routing traffic to a slow pod than to immediately kill it
8. Summary
A readiness probe failure removes your pod from Service endpoints without restarting it. The pod stays Running but receives no traffic. Start by confirming the READY column (not just STATUS) and checking the Events section for Readiness probe failed messages.
| Symptom | Most likely cause | Fix |
|---|---|---|
| READY 0/1, STATUS Running, no restarts | Readiness probe actively failing | Check Events; test endpoint with curl from inside container |
| Fails in first 30–60s only | initialDelaySeconds too low |
Increase delay or add startupProbe |
context deadline exceeded |
timeoutSeconds too low |
Increase to 3–5s; simplify readiness endpoint |
| 503 with dependency error in body | External dependency check in readiness endpoint | Remove dependency checks from /ready; fix or monitor separately |
| Deployment rollout stalls | New pods never reach Ready during rollout | Fix readiness probe; set progressDeadlineSeconds to detect stalls |
| Fails only with Istio enabled | Envoy sidecar not ready before probe fires | Enable holdApplicationUntilProxyStarts in Istio mesh config |
A well-designed readiness probe protects your users from sending requests to pods that can't handle them yet — without causing unnecessary restarts or deployment stalls. Keep the endpoint lightweight, test it under load, and use startupProbe for variable startup times. Note that a pod showing STATUS: Pending rather than Running has a different root cause — check the ImagePullBackOff guide if the pod can't pull its image, or Pod Pending: No Nodes Available if it hasn't been scheduled yet.