
Kubernetes v1.37 introduces a beta, default‑enabled HorizontalPodAutoscaler that can scale workloads down to zero using object or external metrics. Learn how to configure, deploy, and manage zero‑scale workloads safely.
What’s New in v1.37: HPA Scaling to Zero
Kubernetes v1.37 promotes horizontal pod autoscaling (HPA) to zero replicas from an Alpha‑only feature to a Beta API that is enabled by default on the API server and controller manager. This removes the need for external add‑ons or feature‑gate flags to achieve “scale‑to‑zero” behavior for workloads such as queue consumers or batch processors.
Because traditional resource metrics (CPU, memory) are emitted only by running Pods, the HPA cannot decide to scale back up once the replica count reaches zero. v1.37 therefore requires an object or external metric that exists independently of the Pods. A common pattern is to use a queue‑length metric collected by Prometheus.
Typical configuration flow
- Deploy a metrics adapter (e.g., the Prometheus Adapter) that exposes the desired series through the
ExternalMetrics API. - Verify metric availability with a raw API call, for example:
kubectl get --raw '/apis/external.metrics.k8s.io/v1beta1/namespaces/default/queue_consumer_lag?labelSelector=name%3Dworker_tasks' - Create an HPA that references the external metric and sets
minReplicas: 0. A minimal example:apiVersion: autoscaling/v2 kind: HorizontalPodAutoscaler metadata: name: queue-worker spec: scaleTargetRef: apiVersion: apps/v1 kind: Deployment name: queue-worker minReplicas: 0 maxReplicas: 10 metrics: - type: External external: metric: name: queue_consumer_lag selector: matchLabels: name: worker_tasks target: type: Value value: "30" - Deploy the target workload (at least one replica initially). The HPA will reduce the Deployment to zero when the metric reports no pending tasks and will recreate Pods when the metric rises.
v1.37 introduces the ScaledToZero condition on the HPA status. When the controller performs an automatic scale‑down, it records ScaledToZero=True, allowing the controller to continue evaluating external metrics. If an operator manually sets the replica count to zero, the condition remains False, and the HPA treats the workload as paused.
Operational considerations:
- Ensure the downscale stabilization window (default five minutes) matches your workload’s tolerance for brief metric drops.
- During version‑skewed upgrades, confirm that both the API server and controller manager have the feature enabled before creating HPAs with
minReplicas: 0. - When downgrading to a version without the condition‑based implementation, change
minReplicasto at least 1 and scale any zero‑replica workloads back up.
By leveraging object or external metrics, v1.37’s Beta HPA scaling to zero provides a native, declarative mechanism for reducing idle resource consumption while preserving the ability to react to workload demand.
Choosing the Right Metric: Object vs External
CPU and memory metrics are collected by the kubelet from each running pod. When a HorizontalPodAutoscaler (HPA) relies on these resource metrics, the controller must read a value that exists only while at least one replica is active. If the replica count reaches 0, there are no pods from which to obtain CPU or memory data, so the HPA loses its signal and cannot decide to scale back up. This limitation is why traditional resource‑based autoscaling cannot be used to drive a “scale‑to‑zero” workflow.
Object and external metrics solve the problem because they are decoupled from the lifecycle of the pods they control. An object metric (e.g., the length of a Kubernetes Job queue) or an external metric (e.g., a Prometheus series representing pending tasks) continues to be emitted even when no worker pods exist. The HPA can therefore observe a persistent signal and trigger a scale‑up event.
- Independence from pod existence: The metric source lives outside the pod set, so a zero replica count does not silence the metric.
- Predictable scaling logic: The HPA can map a concrete value (e.g., 30 queued tasks) to a desired replica count.
- Compatibility with durable queues: Work that can be buffered (e.g., message queues, task queues) tolerates the cold‑start latency introduced by scaling from zero.
Practical example: a queue consumer deployment named queue‑worker uses the external metric queue_consumer_lag exposed by the Prometheus Adapter. The metric is defined as the sum of pending tasks per queue name:
externalRules:
- seriesQuery: '{__name__="queue_consumer_lag",name!=""}'
metricsQuery: sum(<<.Series>>{<<.LabelMatchers>>}) by (name)
resources:
overrides:
namespace:
resource: namespace
After verifying the metric is reachable with:
kubectl get --raw '/apis/external.metrics.k8s.io/v1beta1/namespaces/default/queue_consumer_lag?labelSelector=name%3Dworker_tasks'
the HPA definition can be created:
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: queue-worker
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: queue-worker
minReplicas: 0
maxReplicas: 10
metrics:
- type: External
external:
metric:
name: queue_consumer_lag
selector:
matchLabels:
name: worker_tasks
target:
type: Value
value: "30"
When the queue is empty, the HPA reduces the deployment to zero replicas. As soon as tasks accumulate, the external metric remains available, the HPA computes the required replica count, and the controller schedules new pods. The ScaledToZero condition recorded by the HPA distinguishes an automatic scale‑down from a manual pause, ensuring that only workloads owned by the HPA are revived.
Step‑by‑Step: Configuring Prometheus Adapter for External Metrics
The Prometheus Adapter facilitates the translation of Prometheus time-series data into metrics compatible with the Kubernetes External Metrics API. By exposing metrics independently of Pod lifecycles, this adapter enables HorizontalPodAutoscaler (HPA) configurations to scale workloads—including those scaled down to zero—based on external signals like queue depth.
Configuring the adapter requires defining externalRules in your configuration manifest. These rules map Prometheus series to the API resources that the HPA controller consumes:
seriesQuery: A Prometheus query pattern used to discover available metrics. It should target the specific metric name while filtering for relevant labels to avoid excessive overhead.metricsQuery: A template that the adapter executes to retrieve the current metric value. It utilizes<<.Series>>and<<.LabelMatchers>>placeholders to map the HPA's requested external metric label selectors to actual Prometheus queries.
A sample configuration snippet for a queue_consumer_lag metric follows:
externalRules:
- seriesQuery: '{__name__="queue_consumer_lag",name!=""}'
metricsQuery: sum(<<.Series>>{<<.LabelMatchers>>}) by (name)
resources:
overrides:
namespace:
resource: namespace
Before deploying an HPA, you must verify that the metrics pipeline is correctly populating the External Metrics API. Use the kubectl get --raw command to query the API directly. This provides immediate feedback on whether the adapter successfully exposes the metric for a specific label selector:
kubectl get --raw '/apis/external.metrics.k8s.io/v1beta1/namespaces/default/queue_consumer_lag?labelSelector=name%3Dworker_tasks'
If this command fails or returns an empty result, the HPA will be unable to calculate replica counts, resulting in ScalingActive=False status. Only proceed to HPA configuration once this query returns a valid JSON payload containing the current metric value, confirming the observability bridge between Prometheus and the Kubernetes control plane is functional.
Defining the HPA: YAML Example and Behavior Details
Starting with Kubernetes v1.37, the HorizontalPodAutoscaler (HPA) natively supports scaling workloads to zero replicas using object or external metrics. This functionality is enabled by default, eliminating the need for custom add-ons. Because CPU and memory metrics are unavailable when no pods are running, this configuration requires metrics that exist independently of the workload, such as queue lengths.
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: queue-worker
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: queue-worker
minReplicas: 0
maxReplicas: 10
metrics:
- type: External
external:
metric:
name: queue_consumer_lag
selector:
matchLabels:
name: worker_tasks
target:
type: Value
value: "30"
The HPA utilizes a ScaledToZero status condition to differentiate between automatic scaling and manual intervention. When the HPA scales a target to zero, it sets ScaledToZero=True, indicating that the controller maintains ownership of the scaling logic and will continue to monitor metrics for scale-up triggers. If a user manually scales a deployment to zero, this condition is not set, and the HPA treats the workload as paused, refusing to perform further scaling actions.
Operational considerations include:
- Stabilization Windows: The default downscale stabilization window remains five minutes. This prevents rapid toggling, though it can be adjusted via
spec.behavior.scaleDown. - Metric Availability: The HPA cannot scale from zero if the metric source is unreachable. If the metrics adapter fails, the HPA will report
ScalingActive=False. - Configuration Constraints: The API server rejects HPA manifests where
minReplicas: 0is combined solely with resource-based metrics (CPU/Memory). - Control Plane Versioning: During upgrades, ensure all control plane components support this feature before implementation, as older controllers may interpret
minReplicas: 0as a manual pause, leaving the workload inactive.
Upgrade, Compatibility, and Operational Considerations
The HPAScaleToZero feature gate is enabled by default on both the kube-apiserver and kube‑controller‑manager starting with Kubernetes v1.37. This default makes the minReplicas: 0 field in an HorizontalPodAutoscaler (HPA) a valid API contract, provided the HPA references an object or external metric.
Version‑skewed upgrade guidance
During a control‑plane upgrade where the API server and controller manager run different versions, follow these steps:
- Confirm that the older component (typically the controller manager) has the
HPAScaleToZerogate enabled; otherwise it will treatreplicas: 0as a manual pause. - Delay creation of HPAs with
minReplicas: 0until both components report the feature as active. - After the upgrade, verify the
ScaledToZerocondition on existing HPAs withkubectl describe hpa <name>.
Preparing to disable the gate or downgrade
Before turning off the feature gate or moving to a version that lacks the condition‑based implementation, perform the following:
- Update every affected HPA definition to set
minReplicasto at least1. - Scale any workload currently at zero to a minimum of one replica (e.g.,
kubectl scale deployment <name> --replicas=1). - Ensure each HPA still has a valid object or external metric; the API server will reject HPAs that rely solely on CPU or memory when
minReplicasis zero. - After the changes, you may safely disable the gate or roll back the cluster version.
Alpha‑to‑Beta evolution
The capability originated as an Alpha feature in Kubernetes v1.16. In v1.36 the ScaledToZero condition was added, allowing the controller to differentiate an automatic scale‑down from a manual pause. v1.37 graduated the feature to Beta and set the gate on by default, completing integration and end‑to‑end testing for external‑metric‑driven scaling to zero.
Practical example
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: queue-worker
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: queue-worker
minReplicas: 0
maxReplicas: 10
metrics:
- type: External
external:
metric:
name: queue_consumer_lag
selector:
matchLabels:
name: worker_tasks
target:
type: Value
value: "30"
Further learning resources
- Official documentation: “Scaling to and from zero”.
- KEP‑2021: “HPA supports scaling to and from zero pods for object and external metrics”.
- Prometheus Adapter guide for exposing external metrics.
- SIG Autoscaling community: Kubernetes Slack
#sig-autoscalingchannel.
Looking for Custom Software or AI Solutions?
Appworks Technologies designs, builds, and scales production enterprise platforms, microservices, and AI agent workflows tailored to your business goals.
Editorial Policy & Research Methodology
Our findings are based on rigorous internal research, verified industry benchmarks, and direct technical implementation experience from our enterprise client projects. All statistics and technical claims are reviewed by senior engineers before publication to ensure accuracy, transparency, and helpfulness for our readers.
