1. Introduction

The Horizontal Pod Autoscaler (HPA) is supposed to be the safety valve of your Kubernetes workloads — automatically adding pods when load increases and removing them when it drops. When it stops working, you're either running with too few replicas during traffic spikes (causing timeouts and degraded performance) or stuck with too many replicas during low traffic (wasting cost).

HPA failures are almost always diagnostic rather than catastrophic — the HPA resource exists and looks correct, but it's not making scaling decisions. The reason is almost always one of: no metrics available, resource requests missing on pods, hitting min/max boundaries, or scale-down stabilisation being too conservative.

2. How HPA Makes Scaling Decisions

The HPA controller polls the Metrics API (provided by the Metrics Server) at --horizontal-pod-autoscaler-sync-period (default 15s). It calculates desired replicas as: ceil(currentReplicas × (currentMetricValue / targetMetricValue)). It then applies minimum, maximum, and stabilisation window constraints before actually changing the replica count.

3. Common Causes

4. Step-by-Step Diagnosis and Fix

Step 1: Read the HPA status

# Check HPA status
kubectl get hpa <hpa-name> -n <namespace>
# TARGETS shows: current/target (e.g. 65%/50%)
# If TARGETS shows <unknown>/50%: metrics problem

# Detailed HPA description
kubectl describe hpa <hpa-name> -n <namespace>
# Look for:
# - "Current replicas" vs "Desired replicas"
# - Conditions section: AbleToScale, ScalingActive, ScalingLimited
# - Events section for scaling decisions and reasons

Step 2: Fix missing or unknown metrics

# "unknown" in HPA TARGETS almost always means:
# 1. Metrics Server is not installed, or
# 2. Pods don't have resource requests set

# Check Metrics Server:
kubectl top pods -n <namespace>
# If this fails, fix Metrics Server first: /fix-kubernetes-metrics-server-not-working.html

# Check pods have resource requests:
kubectl get pod -n <namespace> -l <hpa-selector>   -o jsonpath='{.items[*].spec.containers[*].resources.requests}'

# Fix: add resource requests to the Deployment
spec:
  containers:
  - name: my-app
    resources:
      requests:
        cpu: "100m"      # required for CPU-based HPA
        memory: "128Mi"

Step 3: Check if HPA is hitting limits

# Check current vs min/max
kubectl get hpa <hpa-name> -n <namespace>
# MINPODS and MAXPODS columns

# If REPLICAS = MAXPODS and load is still high:
# The HPA is working — you've hit the ceiling
kubectl patch hpa <hpa-name> -n <namespace>   --type='json' -p='[{"op":"replace","path":"/spec/maxReplicas","value":20}]'

# Describe to see the ScalingLimited condition:
kubectl describe hpa <hpa-name> -n <namespace> | grep -A 3 "ScalingLimited"

Step 4: Check and adjust stabilisation window

# Default scale-down stabilisation is 300s (5 minutes)
# This means HPA won't scale down for 5 minutes after load drops
# This is intentional to prevent flapping

# Adjust if needed:
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: my-app-hpa
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: my-app
  minReplicas: 2
  maxReplicas: 20
  metrics:
  - type: Resource
    resource:
      name: cpu
      target:
        type: Utilization
        averageUtilization: 70
  behavior:
    scaleDown:
      stabilizationWindowSeconds: 60   # reduce from 300s to 60s
    scaleUp:
      stabilizationWindowSeconds: 0    # scale up immediately

Step 5: Test HPA with a load generator

# Generate load to trigger scale-up
kubectl run load-test --image=busybox --rm -it --restart=Never --   sh -c "while true; do wget -q -O- http://<service-name>; done"

# Watch HPA in real time
kubectl get hpa <hpa-name> -n <namespace> -w
# TARGETS should increase; REPLICAS should increase after the sync period

5. Verification Steps

# After fixing, HPA should show non-unknown metrics
kubectl get hpa <hpa-name> -n <namespace>
# TARGETS: 65%/70% (example of real values)

# Under load, replicas should increase
kubectl get pods -n <namespace> -l <selector> -w
# Should see new pods created

# Check HPA events to confirm scaling decisions
kubectl describe hpa <hpa-name> -n <namespace> | grep -A 5 Events

6. Common Mistakes

7. Prevention Tips

8. FAQ

HPA shows the right current CPU but isn't scaling up. Why?

If the current metric is above the target but HPA doesn't scale, check: (1) you're already at maxReplicas, (2) the scale-up stabilisation window hasn't expired yet (check kubectl describe hpa events for "ScalingLimited"), (3) the calculated desired replicas rounds down to current replicas due to the tolerance (HPA ignores changes smaller than 10% by default).

Can I use HPA with memory instead of CPU?

Yes. Use metrics[].resource.name: memory with type: AverageValue (not Utilization — memory percentage isn't useful because memory isn't compressible). For example: target.averageValue: 256Mi. Memory-based HPA is less common because memory doesn't scale down automatically when load drops — pods often retain memory after processing spikes.

9. Summary

HPA behaviourCauseFix
TARGETS shows <unknown>No metrics availableInstall Metrics Server; add resource requests
Doesn't scale up under loadAlready at maxReplicasIncrease maxReplicas; check ScalingLimited condition
Doesn't scale down after loadStabilisation windowReduce stabilizationWindowSeconds in behavior spec
Scales but too slowlyDefault policies too conservativeAdd behavior.scaleUp with shorter stabilisation
Custom metrics unknownPrometheus Adapter missingInstall and configure Prometheus Adapter

Explore More in This Category

Explore more in this category: Kubernetes guides. Browse all DevOps Compass articles or jump to: Kubernetes, AWS, CI/CD, Containers, Monitoring, Networking.