1. Introduction

The Kubernetes Metrics Server is a cluster add-on that collects resource utilisation metrics (CPU and memory) from each node's kubelet and exposes them via the Metrics API. Without it, kubectl top pods and kubectl top nodes don't work, and more critically, the Horizontal Pod Autoscaler can't make scaling decisions. If your HPA is not scaling, a broken Metrics Server is the most likely cause.

This guide covers installation verification, the most common certificate and connectivity errors, and how to configure the Metrics Server correctly for production clusters.

2. What "Metrics Server Not Working" Means

The Metrics Server works by: (1) running as a Deployment in kube-system, (2) connecting to each node's kubelet on port 10250 to scrape resource metrics, (3) aggregating those metrics and registering them with the Kubernetes aggregation layer, and (4) serving them via the metrics.k8s.io API. A failure at any of these steps produces different symptoms.

3. Common Causes

4. Step-by-Step Diagnosis and Fix

Step 1: Confirm the error and check if Metrics Server is installed

# Test kubectl top
kubectl top nodes
# If: "error: Metrics API not available" = not installed or API not registered
# If: "Error from server (ServiceUnavailable)" = installed but not healthy

# Check if Metrics Server is deployed
kubectl get deployment metrics-server -n kube-system
kubectl get pods -n kube-system -l k8s-app=metrics-server

# Check if the metrics API is registered
kubectl api-resources | grep metrics
# Should show: nodes and pods from metrics.k8s.io

Step 2: Install Metrics Server (if missing)

# Official installation
kubectl apply -f https://github.com/kubernetes-sigs/metrics-server/releases/latest/download/components.yaml

# For clusters with self-signed kubelet certificates:
# Download and patch the deployment to skip certificate verification
curl -L https://github.com/kubernetes-sigs/metrics-server/releases/latest/download/components.yaml   -o metrics-server.yaml

# Add --kubelet-insecure-tls to the args in the Deployment
# (Only use this if you understand the security implications)
kubectl patch deployment metrics-server -n kube-system   --type=json   -p='[{"op":"add","path":"/spec/template/spec/containers/0/args/-","value":"--kubelet-insecure-tls"}]'

Step 3: Fix certificate issues (production approach)

# For production, don't use --kubelet-insecure-tls
# Instead, configure kubelet to serve proper certificates

# On each node, check kubelet certificate config:
# /var/lib/kubelet/config.yaml
# serverTLSBootstrap: true   <-- enables automatic cert rotation

# For EKS: kubelet serving certificates are managed automatically
# No --kubelet-insecure-tls needed

# Verify Metrics Server can reach kubelets:
kubectl logs -n kube-system deployment/metrics-server --tail=50
# Look for: "Failed to scrape node"

Step 4: Fix NetworkPolicy issues

# Check if NetworkPolicy blocks Metrics Server → kubelet (port 10250)
kubectl get networkpolicy -n kube-system

# Allow Metrics Server egress to kubelets:
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: allow-metrics-server
  namespace: kube-system
spec:
  podSelector:
    matchLabels:
      k8s-app: metrics-server
  egress:
  - ports:
    - protocol: TCP
      port: 10250
  policyTypes:
  - Egress

Step 5: Fix resource issues for large clusters

# Check if Metrics Server is OOMKilled
kubectl describe pod -n kube-system -l k8s-app=metrics-server | grep -A 5 "Last State"

# Increase Metrics Server resources for large clusters:
kubectl patch deployment metrics-server -n kube-system   --type=json   -p='[{"op":"replace","path":"/spec/template/spec/containers/0/resources",
       "value":{"requests":{"cpu":"200m","memory":"200Mi"},
                "limits":{"cpu":"1","memory":"512Mi"}}}]'

5. Verification Steps

# These should all return data after fixing
kubectl top nodes
kubectl top pods -A --sort-by=memory | head -10

# Confirm the metrics API is working
kubectl get --raw /apis/metrics.k8s.io/v1beta1/nodes | jq '.items[].metadata.name'

# HPA should now have access to metrics
kubectl describe hpa <hpa-name> -n <namespace>
# Look for: "AbleToScale: True" and recent metric values

6. Common Mistakes

7. Prevention Tips

8. FAQ

kubectl top shows some nodes but not others. Why?

The Metrics Server couldn't scrape certain kubelets. Check the Metrics Server logs for "failed to scrape node" errors with the specific node names. Causes: the node's kubelet isn't reachable from the Metrics Server pod, the node's certificate is failing verification, or there's a NetworkPolicy blocking the scrape.

Metrics Server is running but HPA says "unable to get metrics". Why?

The Metrics API is available but returning no data for the HPA's target resource. Usually this means: (1) the HPA references CPU/memory metrics but the pod hasn't run long enough to have metrics, (2) the targeted pods don't have resource requests set (required for percentage-based CPU metrics), or (3) the Metrics Server has a scrape delay and HPA is reading stale data. See Fix HPA Not Scaling for the full diagnosis.

9. Summary

ErrorCauseFix
Metrics API not availableNot installedInstall via official manifests
ServiceUnavailablePod crashing or not readyCheck pod logs; fix cert or network issues
Failed to scrape nodeCan't reach kubelet on 10250Check NetworkPolicy; fix cert verification
OOMKilledNot enough memory for cluster sizeIncrease Metrics Server memory limits
kubectl top no dataMetrics API registered but emptyWait for first scrape; check kubelet health

Explore More in This Category

Explore more in this category: Kubernetes guides. Browse all DevOps Compass articles or jump to: Kubernetes, AWS, CI/CD, Containers, Monitoring, Networking.