If your Kubernetes HPA scales down slowly, it is almost certainly doing exactly what it was designed to do. By default the HorizontalPodAutoscaler waits for a 300-second stabilization window before removing pods: it keeps the highest replica recommendation it has computed in the last five minutes, and only drops below that once every recommendation in the window agrees. Add the 15-second sync period, the metrics-server scrape interval and a 10% tolerance band, and “traffic stopped at 10:00, pods started going away at 10:06” is the normal, healthy outcome.
The fix, if you want a different outcome, is the behavior field in autoscaling/v2. This is the short version for “scale down faster”:
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: api
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: api
minReplicas: 2
maxReplicas: 30
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70
behavior:
scaleDown:
stabilizationWindowSeconds: 60 # default is 300
policies:
- type: Percent
value: 50
periodSeconds: 30The rest of this article explains every moving part of HPA scale-down behavior — so you can tell the difference between an HPA that is slow on purpose, one that is stuck, and one that is flapping — and gives tested recipes for each case. If your HPA runs on memory, the memory-specific details (why memory does not fall after load drops, requests vs limits) are covered in Kubernetes HPA on memory; this page is about the scale-down machinery itself, which is the same for every metric.
What Happens Between “Load Dropped” and “Pod Removed”
When people say the HPA is slow to scale down, they usually mean the total elapsed time from the moment load falls to the moment a pod is terminated. That time is the sum of several independent delays, and it helps to see them in order:
- Metric collection. metrics-server scrapes the kubelets every 15 seconds by default, and the kubelet’s CPU figure is itself a rate over a short window. A drop in load shows up in the resource metrics API somewhere between a few seconds and ~30 seconds later.
- HPA sync period. The HPA controller in
kube-controller-managerevaluates every HPA every 15 seconds (--horizontal-pod-autoscaler-sync-period). Add up to another 15 seconds. - Tolerance. If the ratio between the current metric and the target is within 10% of 1.0, the controller does nothing. A gradual decline can spend minutes inside that band.
- Stabilization window. The controller records a recommendation every sync and, for scale-down, applies the highest recommendation from the last 300 seconds. This is the big one: five minutes minimum after the last high recommendation.
- Scaling policies. Once the window allows it, the default scale-down policy removes up to 100% of the excess pods every 15 seconds — so this step is fast by default, unless you slowed it down.
- Pod termination. The Deployment controller deletes pods, which then go through
preStophooks andterminationGracePeriodSeconds. A 60-second graceful shutdown is 60 more seconds before the pod is actually gone.
So with defaults, a clean drop from 10 replicas to 3 typically completes between 5.5 and 6.5 minutes after load falls. If you are seeing 20 minutes, or never, something else is going on — keep reading.
The Default HPA Behavior (and Why Scale-Down Is Deliberately Slow)
When you omit behavior, the HPA behaves as if you had written this (the values are straight from the Kubernetes documentation):
behavior:
scaleDown:
stabilizationWindowSeconds: 300
policies:
- type: Percent
value: 100
periodSeconds: 15
scaleUp:
stabilizationWindowSeconds: 0
policies:
- type: Percent
value: 100
periodSeconds: 15
- type: Pods
value: 4
periodSeconds: 15
selectPolicy: MaxThe asymmetry is intentional. Scaling up late costs you latency and errors; scaling down late costs you a few minutes of spare capacity. So scale-up reacts immediately (no window, double the fleet or add four pods every 15 seconds, whichever is larger) and scale-down waits five minutes to make sure the drop is real.
The cluster-wide default for the scale-down window comes from the --horizontal-pod-autoscaler-downscale-stabilization flag on kube-controller-manager (5 minutes). On managed platforms — EKS, GKE, AKS — you cannot change controller-manager flags, which is exactly why the per-HPA behavior field exists. Always tune in the HPA object, never rely on the cluster flag.
How stabilizationWindowSeconds Actually Works
The stabilization window is the most misunderstood field in the HPA spec, because the name suggests a delay or a cooldown. It is neither. It is a rolling maximum (for scale-down) or rolling minimum (for scale-up) over the replica recommendations computed during the window.
Concretely: every 15 seconds the controller computes a desired replica count from the metrics and stores it with a timestamp. For scale-down, before acting it looks at all recommendations from the last stabilizationWindowSeconds and uses the highest one. A single high recommendation four minutes ago is enough to hold the fleet at that size for one more minute.
Walk through an example with the default 300-second window:
| Time | Load | Recommendation | Recommendations in window | Applied |
|---|---|---|---|---|
| 10:00 | high | 10 | 10 | 10 |
| 10:01 | low | 4 | 10, 4 | 10 |
| 10:03 | low | 3 | 10, 4, 3 | 10 |
| 10:05 | low | 3 | 10 (at 10:00), 4, 3 | 10 |
| 10:05:15 | low | 3 | 4, 3, 3 | 4 |
| 10:06:15 | low | 3 | 3, 3 | 3 |
This is why the scale-down happens in steps that mirror the recommendations from five minutes earlier, and why a brief spike during the quiet period resets the clock. While the window is holding replicas up, kubectl describe hpa shows the condition AbleToScale True ScaleDownStabilized with the message “recent recommendations were higher than current one, applying the highest recent recommendation”. If you see that message, the HPA is not broken — it is waiting.
A few properties of the field worth knowing:
- The value range is 0 to 3600 seconds. Anything above one hour is rejected by API validation.
stabilizationWindowSeconds: 0on scale-down means “act on the current recommendation immediately”. Useful for batch workers; dangerous for anything with bursty traffic.- The same field exists under
scaleUp, where it works as a rolling minimum. Setting a scale-up window of 60 seconds means the HPA only scales up if the last minute of recommendations all agree — a good way to ignore one-sample spikes, at the cost of reacting a minute later. - The window is held in the controller’s memory. When
kube-controller-managerrestarts or fails over, the recommendation history is lost, and the next scale-down can happen sooner than you expect.
Scaling Policies: Pods, Percent, periodSeconds and selectPolicy
Policies limit how fast the replica count can change once the stabilization window has allowed a change. Each policy says “in any periodSeconds window, change by at most value pods (or value percent of the pods)”.
type: Pods is an absolute number: value: 2 with periodSeconds: 60 means at most two pods removed per minute, regardless of fleet size. This is the predictable option and the one I default to for scale-down.
type: Percent is relative to the replica count at the start of the period: value: 25 with periodSeconds: 60 on a 40-pod fleet allows removing 10 pods in that minute, but on a 4-pod fleet only 1. Percent policies scale with the fleet, which is what you want for large deployments and a nuisance for small ones.
periodSeconds can go up to 1800 (30 minutes). The controller tracks actual scaling events within the period, so a policy of 1 pod per 300 seconds really means “no more than one removal in any rolling five-minute span”, even across multiple HPA syncs.
When you list several policies, selectPolicy decides which one wins:
Max(the default) picks the policy that allows the biggest change. WithPods: 4andPercent: 100, a 2-pod fleet can grow by 4 and a 20-pod fleet by 20.Minpicks the policy that allows the smallest change. This is the conservative choice for scale-down: “at most 10% or 2 pods per minute, whichever is smaller”.Disabledturns scaling off in that direction completely.scaleDown.selectPolicy: Disabledmeans the HPA will only ever add pods.
When a policy is what’s holding the fleet back, kubectl describe hpa shows ScalingLimited True ScaleDownLimit — “the desired replica count is decreasing faster than the maximum scale rate”. That is different from ScaleDownStabilized, and it tells you which knob to turn: the window or the policy.
Tolerance and Rounding: Why the HPA Sometimes Never Scales Down
A frequent complaint is not “the HPA is slow”, but “the HPA never scales down, even though usage is clearly below target”. Two pieces of arithmetic cause most of these cases.
The 10% tolerance band. The controller computes ratio = currentMetricValue / targetValue and skips any action if the ratio is within ±0.1 of 1.0. With a 70% CPU target, an average of 64% gives a ratio of 0.914 — inside the band, so nothing happens. Your fleet can sit at 64% indefinitely, 6 points below target, without a single pod removed.
Ceiling rounding at small replica counts. The desired count is ceil(currentReplicas × ratio). Rounding up is safe for scale-up, but it makes scale-down surprisingly sticky on small fleets. With 3 replicas, a 70% target and an average of 50%:
desired = ceil(3 × 50 / 70) = ceil(2.14) = 3The ratio (0.71) is well outside the tolerance, yet the result is still 3. To go from 3 to 2 pods, the average must drop to 46.7% or below (2/3 × 70), because those 2 pods would then run at 70%. The HPA is refusing to remove a pod that would push the remaining ones over target — which is correct, but it looks like a bug when you are staring at a dashboard that says 50%. The smaller the fleet, the bigger this effect: going from 2 to 1 requires the average to fall to half the target.
Configurable tolerance per HPA. Until recently the 10% was a cluster-wide constant (--horizontal-pod-autoscaler-tolerance), which on managed clusters meant you could not change it at all. The HPAConfigurableTolerance feature adds a tolerance field per direction:
behavior:
scaleUp:
tolerance: 0.05 # react when 5% above target
scaleDown:
tolerance: 0.15 # only scale down when 15% below targetThe feature went alpha in Kubernetes 1.33, beta in 1.35 (still disabled by default in beta, so managed providers generally did not expose it) and GA in 1.37, where it is always on. On a cluster older than 1.37, check the feature gate before relying on it — an API server that does not know the field silently drops it. A lower scale-up tolerance with a higher scale-down tolerance is a sensible pairing for latency-sensitive services: react early to growth, ignore small declines.
Why an HPA Stops Scaling Down: The Checklist
When the HPA is not scaling down at all — not slowly, not at all — work through these in order. They cover the cases I have actually seen in production.
minReplicas is reached. Obvious, but check it first. ScalingLimited True TooFewReplicas in the conditions confirms it.
Another metric is holding it up. With multiple metrics, the HPA computes a desired count for each and uses the largest. CPU can be at 10% while memory or a queue-length metric keeps the fleet at its current size. kubectl describe hpa lists each metric’s current value; look for the one still near its target. For memory specifically, the reasons it does not fall after load drops are covered in the HPA memory guide.
One metric is failing. This one is subtle and documented: if any metric cannot be fetched and the remaining metrics suggest a scale-down, the controller skips scaling entirely. Scale-up still works in that situation; scale-down does not. A broken Prometheus Adapter query or a missing custom metric can therefore freeze your fleet at its peak size for days. Look for FailedGetPodsMetric, FailedGetExternalMetric or FailedGetResourceMetric events.
Pods with missing metrics. When some pods have no metrics yet (new pods, or pods on a node where metrics-server is failing), the controller recomputes conservatively and assumes those pods are at 100% of target when evaluating a scale-down. A handful of pods with no metrics can keep the average above the scale-down threshold.
Rounding on a small fleet. See the previous section. On 2 or 3 replicas, the average may need to fall far below the target before a pod is removed.
Scale-down is disabled. Someone may have set scaleDown.selectPolicy: Disabled, or a very long window. kubectl get hpa <name> -o yaml and read the behavior block — including defaults, which the API server fills in.
Something else owns the replica count. A GitOps tool that sets spec.replicas on the Deployment, a KEDA ScaledObject targeting the same Deployment, or a second HPA will fight the HPA. ArgoCD will even show the Deployment as out of sync every time the HPA changes it. Remove replicas from the manifest you apply, and make sure exactly one autoscaler targets each workload.
Requests are wrong. Utilization is usage divided by request. If requests are set far below real idle usage, “idle” is already above target and the HPA will never scale down. The fix is to right-size requests, not to tune the HPA — see Kubernetes resource requests and limits.
Why an HPA Scales Down Too Fast (or Flaps)
The opposite problem shows up on bursty workloads: the HPA removes pods, traffic comes back two minutes later, the HPA adds them again, and new pods take a while to be useful. Every cycle costs cold starts, connection re-balancing and, often, a latency spike.
Common causes:
- A short or zero scale-down window. Teams set
stabilizationWindowSeconds: 0to “fix” slowness and create flapping. If traffic has a periodicity shorter than your window — a cron that fires every 10 minutes, a batch that arrives every quarter hour — the window must be longer than that period. - Aggressive Percent policies.
Percent: 100allows removing all excess pods at once. APodspolicy that removes one or two per period turns a cliff into a staircase, which gives you time to notice a returning load. - Slow-starting pods. If a new pod takes three minutes to warm up — JVM JIT compilation, cache loading, connection pools — every scale-down you later have to reverse costs three minutes of under-capacity. The HPA has two cluster-level safeguards for CPU (
--horizontal-pod-autoscaler-cpu-initialization-period, 5 minutes, and--horizontal-pod-autoscaler-initial-readiness-delay, 30 seconds) that ignore CPU samples from pods that are still starting. They only help if your readiness probe does not report Ready before the warm-up is done. Use astartupProbethat holds the pod back until it can actually serve.
HPA Behavior Recipes (autoscaling/v2)
All of these are complete manifests you can apply as-is after changing the target and metrics. They use CPU for brevity; the behavior block is the part that matters, and it works identically with memory, custom or external metrics.
Scale down faster
For stateless services with fast startup and load that drops cleanly — internal APIs, dev and staging environments, anything where idle capacity costs more than an occasional cold start:
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: api-fast-down
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: api
minReplicas: 2
maxReplicas: 30
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70
behavior:
scaleDown:
stabilizationWindowSeconds: 60
policies:
- type: Percent
value: 50
periodSeconds: 30
selectPolicy: MaxThe fleet can shrink one minute after load drops, halving at most every 30 seconds. Going below 60 seconds rarely buys anything real: metric lag already eats 15-30 seconds of it.
Scale down gradually (fast up, slow down)
The safe default for customer-facing services, bursty traffic and slow-starting runtimes:
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: api-gradual
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: api
minReplicas: 3
maxReplicas: 50
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 65
behavior:
scaleUp:
stabilizationWindowSeconds: 0
policies:
- type: Percent
value: 100
periodSeconds: 30
- type: Pods
value: 4
periodSeconds: 30
selectPolicy: Max
scaleDown:
stabilizationWindowSeconds: 600
policies:
- type: Pods
value: 2
periodSeconds: 120
- type: Percent
value: 10
periodSeconds: 120
selectPolicy: MinScale-up doubles the fleet (or adds 4, whichever is more) every 30 seconds. Scale-down waits until ten minutes of recommendations agree, then removes at most 2 pods or 10% every two minutes, whichever is smaller. A drop from 50 to 10 pods takes about 40 minutes, which is exactly the point: if traffic comes back, you are still mostly provisioned.
Never scale down automatically
For workloads where the HPA is there to absorb peaks, and scale-down should happen only through a deploy or a human decision — stateful consumers that rebalance partitions on every change, services with very expensive warm-up, or the first weeks of a new HPA you do not yet trust:
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: consumer-no-down
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: consumer
minReplicas: 4
maxReplicas: 24
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70
behavior:
scaleDown:
selectPolicy: DisabledThe cost: the fleet sits at its high-water mark until someone lowers maxReplicas or redeploys.
KEDA: cooldownPeriod Is Not What Controls Scale-Down
If you use KEDA, you have probably tried cooldownPeriod to slow down or speed up scale-down and seen no effect. That is because cooldownPeriod (default 300 seconds) only applies to the transition from 1 replica to 0: how long KEDA waits after the last active trigger before scaling the workload to zero. Every change between N and 1 replicas is made by the HPA that KEDA creates, and is controlled by the same behavior field described above — passed through the ScaledObject:
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
name: worker
spec:
scaleTargetRef:
name: worker
minReplicaCount: 0
maxReplicaCount: 40
cooldownPeriod: 300 # only 1 -> 0
advanced:
horizontalPodAutoscalerConfig:
behavior: # everything between maxReplicaCount and 1
scaleDown:
stabilizationWindowSeconds: 120
policies:
- type: Percent
value: 25
periodSeconds: 60
triggers:
- type: rabbitmq
metadata:
queueName: jobs
mode: QueueLength
value: "20"
authenticationRef:
name: rabbitmq-authFor queue-driven workers, a short window is usually fine: queue length is a leading signal, not a lagging one like CPU. More on KEDA triggers in event-driven autoscaling with KEDA.
How to Diagnose HPA Scale-Down with kubectl
Three commands tell you almost everything.
Watch the HPA live to see the relationship between the metric and the replica count over time:
kubectl get hpa api -wDescribe it to see the conditions and the recent events — this is where the “why” lives:
kubectl describe hpa apiA healthy HPA waiting on its stabilization window looks like this:
Metrics: ( current / target )
resource cpu on pods (as a percentage of request): 18% (90m) / 70%
Min replicas: 2
Max replicas: 30
Behavior:
Scale Up:
Stabilization Window: 0 seconds
Select Policy: Max
Policies:
- Type: Pods Value: 4 Period: 15 seconds
- Type: Percent Value: 100 Period: 15 seconds
Scale Down:
Stabilization Window: 300 seconds
Select Policy: Max
Policies:
- Type: Percent Value: 100 Period: 15 seconds
Deployment pods: 10 current / 10 desired
Conditions:
Type Status Reason Message
---- ------ ------ -------
AbleToScale True ScaleDownStabilized recent recommendations were higher than current one, applying the highest recent recommendation
ScalingActive True ValidMetricFound the HPA was able to successfully calculate a replica count from cpu resource utilization (percentage of request)
ScalingLimited False DesiredWithinRange the desired count is within the acceptable rangeHow to read the conditions:
| Condition | Reason | Meaning |
|---|---|---|
| AbleToScale | ScaleDownStabilized | The stabilization window is holding replicas up. Wait, or shorten the window. |
| AbleToScale | ReadyForNewScale | No stabilization in effect; the recommendation is being applied as-is. |
| ScalingLimited | ScaleDownLimit | A scale-down policy is capping the rate. Loosen the policy if too slow. |
| ScalingLimited | TooFewReplicas | minReplicas reached. |
| ScalingLimited | DesiredWithinRange | No limit is being applied. |
| ScalingActive | ValidMetricFound | Metrics are OK. Anything else here means the HPA cannot calculate at all. |
If ScalingActive is False, stop tuning behavior: the HPA is not receiving metrics, and the events at the bottom of the describe output (FailedGetResourceMetric, FailedComputeMetricsReplicas) say why.
Check the raw metrics the HPA sees, to rule out metrics-server lag or missing pods:
kubectl get --raw "/apis/metrics.k8s.io/v1beta1/namespaces/default/pods" | jq '.items[] | {name: .metadata.name, cpu: .containers[].usage.cpu}'
kubectl top pods -l app=apiIf some pods are missing from that output, you have found your “missing metrics treated as 100%” problem.
Pod Scale-Down Is Not Node Scale-Down
One last source of confusion: even when the HPA removes pods promptly, the nodes do not disappear at the same time, so the cloud bill does not drop when you expect. Node removal is a separate decision made by Cluster Autoscaler or Karpenter, with its own delays (Cluster Autoscaler waits 10 minutes by default before removing an underutilized node) and its own blockers, such as PodDisruptionBudgets and pods that cannot be moved. If cost is the reason you want faster scale-down, tune both layers — the comparison of how each one consolidates nodes is in Cluster Autoscaler vs Karpenter.
And before tuning behavior at all, make sure the metric is the right one: a CPU target on a service that is really bound by I/O or a downstream dependency will never scale down cleanly, no matter what the window is. The broader design questions — which metric, which target, when HPA is the wrong tool — are in Kubernetes HPA best practices.
Frequently Asked Questions
Why does my Kubernetes HPA scale down so slowly?
Because the default scale-down stabilization window is 300 seconds: the HPA applies the highest replica recommendation from the last five minutes, so pods are only removed once five minutes of recommendations agree that fewer are needed. Add 15-30 seconds of metric lag, the 15-second HPA sync period and pod termination time, and a scale-down typically starts 5.5 to 6.5 minutes after load drops. Set spec.behavior.scaleDown.stabilizationWindowSeconds to a lower value (for example 60) to make it faster.
What is the default stabilizationWindowSeconds for HPA?
For scale-down it is 300 seconds, taken from the --horizontal-pod-autoscaler-downscale-stabilization flag of kube-controller-manager unless you set it in the HPA’s behavior field. For scale-up it is 0 seconds, which means the HPA scales up as soon as a higher recommendation is computed. The field accepts values from 0 to 3600 seconds.
How do I make the HPA scale down faster?
Add a behavior.scaleDown block to your autoscaling/v2 HPA with a shorter stabilizationWindowSeconds (60 is a reasonable floor) and a permissive policy such as type: Percent, value: 50, periodSeconds: 30. If it still does not scale down, the window is not the problem: check for another metric holding the fleet up, a metric that fails to fetch, minReplicas, or the ceiling rounding that makes small fleets sticky.
Why is my HPA not scaling down even though CPU is below the target?
The most common reasons are the 10% tolerance band (at a 70% target, an average of 64% triggers nothing), ceiling rounding on small fleets (3 pods at 50% against a 70% target still compute to 3), a second metric that is still near its target, or a metric that cannot be fetched — in which case the HPA skips scale-down entirely. kubectl describe hpa shows each metric’s value, the conditions and the events that tell you which one applies.
What is the difference between stabilizationWindowSeconds and periodSeconds?
stabilizationWindowSeconds decides whether to scale: it takes the highest (scale-down) or lowest (scale-up) recommendation in the window, which filters out short-lived fluctuations. periodSeconds belongs to a scaling policy and decides how fast to scale once a change is allowed: at most value pods or percent within any periodSeconds span. The window shows up as ScaleDownStabilized in the HPA conditions, the policy as ScaleDownLimit.
Can I change the HPA tolerance per HorizontalPodAutoscaler?
Yes, from Kubernetes 1.37 the tolerance field under behavior.scaleUp and behavior.scaleDown is GA and always available, so you can set, for example, 5% for scale-up and 15% for scale-down on a single HPA. It was alpha in 1.33 and beta (disabled by default) in 1.35 and 1.36, so on older clusters it depends on the feature gate. Without it, every HPA uses the cluster-wide 10% from --horizontal-pod-autoscaler-tolerance.
Does KEDA cooldownPeriod control HPA scale-down?
No. KEDA’s cooldownPeriod (default 300 seconds) only controls how long KEDA waits before scaling from 1 replica to 0. Scaling between the maximum and 1 replica is done by the HPA that KEDA creates, and you tune it with spec.advanced.horizontalPodAutoscalerConfig.behavior in the ScaledObject, using the same scaleDown window and policies as a plain HPA.
How do I stop the HPA from scaling down at all?
Set spec.behavior.scaleDown.selectPolicy: Disabled. The HPA will still scale up when metrics exceed the target, but it will never remove pods on its own; the replica count stays at its high-water mark until you lower maxReplicas, recreate the HPA or scale down during a deploy.
Conclusion
A Kubernetes HPA that scales down slowly is, nine times out of ten, the default 300-second stabilization window doing its job, plus metric and sync lag on top. Treat that as a starting point, not a bug: shorten the window for services that start fast and cost money when idle, lengthen it and add a Pods policy for services where every cold start hurts, and disable scale-down entirely where only a human should make that call.
When the HPA does not scale down at all, stop tuning behavior and read kubectl describe hpa. The conditions tell you whether you are looking at stabilization (ScaleDownStabilized), a rate limit (ScaleDownLimit), a floor (TooFewReplicas) or broken metrics (ScalingActive False) — and the tolerance and rounding arithmetic explains the rest. Get that right, and the HPA becomes predictable enough that you can reason about its next move instead of waiting for it.