1. Introduction
The Horizontal Pod Autoscaler (HPA) is supposed to be the safety valve of your Kubernetes workloads — automatically adding pods when load increases and removing them when it drops. When it stops working, you're either running with too few replicas during traffic spikes (causing timeouts and degraded performance) or stuck with too many replicas during low traffic (wasting cost).
HPA failures are almost always diagnostic rather than catastrophic — the HPA resource exists and looks correct, but it's not making scaling decisions. The reason is almost always one of: no metrics available, resource requests missing on pods, hitting min/max boundaries, or scale-down stabilisation being too conservative.
2. How HPA Makes Scaling Decisions
The HPA controller polls the Metrics API (provided by the Metrics Server) at --horizontal-pod-autoscaler-sync-period (default 15s). It calculates desired replicas as: ceil(currentReplicas × (currentMetricValue / targetMetricValue)). It then applies minimum, maximum, and stabilisation window constraints before actually changing the replica count.
3. Common Causes
- Metrics Server is not installed — HPA can't get any metrics
- Pods don't have
resources.requests.cpuset — CPU percentage HPA shows "unknown" - Already at
maxReplicas— can't scale up further - Already at
minReplicas— can't scale down further - Scale-down stabilisation window — HPA waits 5 minutes before scaling down by default
- Metrics are available but below the target threshold — HPA decided no change is needed
- For custom metrics: Prometheus Adapter not installed or misconfigured
- The Deployment's selector doesn't match the pods actually running
4. Step-by-Step Diagnosis and Fix
Step 1: Read the HPA status
# Check HPA status
kubectl get hpa <hpa-name> -n <namespace>
# TARGETS shows: current/target (e.g. 65%/50%)
# If TARGETS shows <unknown>/50%: metrics problem
# Detailed HPA description
kubectl describe hpa <hpa-name> -n <namespace>
# Look for:
# - "Current replicas" vs "Desired replicas"
# - Conditions section: AbleToScale, ScalingActive, ScalingLimited
# - Events section for scaling decisions and reasons
Step 2: Fix missing or unknown metrics
# "unknown" in HPA TARGETS almost always means:
# 1. Metrics Server is not installed, or
# 2. Pods don't have resource requests set
# Check Metrics Server:
kubectl top pods -n <namespace>
# If this fails, fix Metrics Server first: /fix-kubernetes-metrics-server-not-working.html
# Check pods have resource requests:
kubectl get pod -n <namespace> -l <hpa-selector> -o jsonpath='{.items[*].spec.containers[*].resources.requests}'
# Fix: add resource requests to the Deployment
spec:
containers:
- name: my-app
resources:
requests:
cpu: "100m" # required for CPU-based HPA
memory: "128Mi"
Step 3: Check if HPA is hitting limits
# Check current vs min/max
kubectl get hpa <hpa-name> -n <namespace>
# MINPODS and MAXPODS columns
# If REPLICAS = MAXPODS and load is still high:
# The HPA is working — you've hit the ceiling
kubectl patch hpa <hpa-name> -n <namespace> --type='json' -p='[{"op":"replace","path":"/spec/maxReplicas","value":20}]'
# Describe to see the ScalingLimited condition:
kubectl describe hpa <hpa-name> -n <namespace> | grep -A 3 "ScalingLimited"
Step 4: Check and adjust stabilisation window
# Default scale-down stabilisation is 300s (5 minutes)
# This means HPA won't scale down for 5 minutes after load drops
# This is intentional to prevent flapping
# Adjust if needed:
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: my-app-hpa
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: my-app
minReplicas: 2
maxReplicas: 20
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70
behavior:
scaleDown:
stabilizationWindowSeconds: 60 # reduce from 300s to 60s
scaleUp:
stabilizationWindowSeconds: 0 # scale up immediately
Step 5: Test HPA with a load generator
# Generate load to trigger scale-up
kubectl run load-test --image=busybox --rm -it --restart=Never -- sh -c "while true; do wget -q -O- http://<service-name>; done"
# Watch HPA in real time
kubectl get hpa <hpa-name> -n <namespace> -w
# TARGETS should increase; REPLICAS should increase after the sync period
5. Verification Steps
# After fixing, HPA should show non-unknown metrics
kubectl get hpa <hpa-name> -n <namespace>
# TARGETS: 65%/70% (example of real values)
# Under load, replicas should increase
kubectl get pods -n <namespace> -l <selector> -w
# Should see new pods created
# Check HPA events to confirm scaling decisions
kubectl describe hpa <hpa-name> -n <namespace> | grep -A 5 Events
6. Common Mistakes
- Using CPU percentage HPA without setting
requests.cpu— you'll always see "unknown" metrics - Setting
maxReplicastoo low for expected traffic — HPA can't help if it's already at the ceiling - Expecting immediate scale-down — the 5-minute stabilisation window is intentional and important
- Not testing HPA before going to production — always validate with a load test in staging first
- Using HPA with a StatefulSet that has PVCs — each replica needs its own PVC, which may not provision fast enough
7. Prevention Tips
- Always set
resources.requestson pods targeted by HPA — it's required for CPU percentage metrics - Set sensible
minReplicas(at least 2 for production) to ensure availability during cold starts - Define explicit
behaviorblocks in HPA to tune scale-up and scale-down speeds for your workload - Monitor HPA events and the
ScalingLimitedcondition — hitting maxReplicas silently without alerting is a common issue - Ensure Metrics Server is healthy — HPA is completely disabled without it
- Use
kubectl top podsto regularly validate that metrics are flowing before incidents
8. FAQ
HPA shows the right current CPU but isn't scaling up. Why?
If the current metric is above the target but HPA doesn't scale, check: (1) you're already at maxReplicas, (2) the scale-up stabilisation window hasn't expired yet (check kubectl describe hpa events for "ScalingLimited"), (3) the calculated desired replicas rounds down to current replicas due to the tolerance (HPA ignores changes smaller than 10% by default).
Can I use HPA with memory instead of CPU?
Yes. Use metrics[].resource.name: memory with type: AverageValue (not Utilization — memory percentage isn't useful because memory isn't compressible). For example: target.averageValue: 256Mi. Memory-based HPA is less common because memory doesn't scale down automatically when load drops — pods often retain memory after processing spikes.
9. Summary
| HPA behaviour | Cause | Fix |
|---|---|---|
| TARGETS shows <unknown> | No metrics available | Install Metrics Server; add resource requests |
| Doesn't scale up under load | Already at maxReplicas | Increase maxReplicas; check ScalingLimited condition |
| Doesn't scale down after load | Stabilisation window | Reduce stabilizationWindowSeconds in behavior spec |
| Scales but too slowly | Default policies too conservative | Add behavior.scaleUp with shorter stabilisation |
| Custom metrics unknown | Prometheus Adapter missing | Install and configure Prometheus Adapter |
Explore More in This Category
Explore more in this category: Kubernetes guides. Browse all DevOps Compass articles or jump to: Kubernetes, AWS, CI/CD, Containers, Monitoring, Networking.