Articles

Kubernetes v1.37: Scale Workloads to Zero with HorizontalPodAutoscaler

Kubernetes v1.37 introduces a beta, default‑enabled HorizontalPodAutoscaler that can scale workloads down to zero using object or external metrics. Learn how to configure, deploy, and manage zero‑scale workloads safely.

Written by:
APin

Senior Technology Analyst • Verified Expert

More from this author →
Kubernetes v1.37: Scale Workloads to Zero with HorizontalPodAutoscaler

Kubernetes v1.37 introduces a beta, default‑enabled HorizontalPodAutoscaler that can scale workloads down to zero using object or external metrics. Learn how to configure, deploy, and manage zero‑scale workloads safely.

What’s New in v1.37: HPA Scaling to Zero

Kubernetes v1.37 promotes horizontal pod autoscaling (HPA) to zero replicas from an Alpha‑only feature to a Beta API that is enabled by default on the API server and controller manager. This removes the need for external add‑ons or feature‑gate flags to achieve “scale‑to‑zero” behavior for workloads such as queue consumers or batch processors.

Because traditional resource metrics (CPU, memory) are emitted only by running Pods, the HPA cannot decide to scale back up once the replica count reaches zero. v1.37 therefore requires an object or external metric that exists independently of the Pods. A common pattern is to use a queue‑length metric collected by Prometheus.

Typical configuration flow

  1. Deploy a metrics adapter (e.g., the Prometheus Adapter) that exposes the desired series through the ExternalMetrics API.
  2. Verify metric availability with a raw API call, for example:
    kubectl get --raw '/apis/external.metrics.k8s.io/v1beta1/namespaces/default/queue_consumer_lag?labelSelector=name%3Dworker_tasks'
  3. Create an HPA that references the external metric and sets minReplicas: 0. A minimal example:
    apiVersion: autoscaling/v2
    kind: HorizontalPodAutoscaler
    metadata:
      name: queue-worker
    spec:
      scaleTargetRef:
        apiVersion: apps/v1
        kind: Deployment
        name: queue-worker
      minReplicas: 0
      maxReplicas: 10
      metrics:
      - type: External
        external:
          metric:
            name: queue_consumer_lag
            selector:
              matchLabels:
                name: worker_tasks
          target:
            type: Value
            value: "30"
  4. Deploy the target workload (at least one replica initially). The HPA will reduce the Deployment to zero when the metric reports no pending tasks and will recreate Pods when the metric rises.

v1.37 introduces the ScaledToZero condition on the HPA status. When the controller performs an automatic scale‑down, it records ScaledToZero=True, allowing the controller to continue evaluating external metrics. If an operator manually sets the replica count to zero, the condition remains False, and the HPA treats the workload as paused.

Operational considerations:

  • Ensure the downscale stabilization window (default five minutes) matches your workload’s tolerance for brief metric drops.
  • During version‑skewed upgrades, confirm that both the API server and controller manager have the feature enabled before creating HPAs with minReplicas: 0.
  • When downgrading to a version without the condition‑based implementation, change minReplicas to at least 1 and scale any zero‑replica workloads back up.

By leveraging object or external metrics, v1.37’s Beta HPA scaling to zero provides a native, declarative mechanism for reducing idle resource consumption while preserving the ability to react to workload demand.

Choosing the Right Metric: Object vs External

CPU and memory metrics are collected by the kubelet from each running pod. When a HorizontalPodAutoscaler (HPA) relies on these resource metrics, the controller must read a value that exists only while at least one replica is active. If the replica count reaches 0, there are no pods from which to obtain CPU or memory data, so the HPA loses its signal and cannot decide to scale back up. This limitation is why traditional resource‑based autoscaling cannot be used to drive a “scale‑to‑zero” workflow.

Object and external metrics solve the problem because they are decoupled from the lifecycle of the pods they control. An object metric (e.g., the length of a Kubernetes Job queue) or an external metric (e.g., a Prometheus series representing pending tasks) continues to be emitted even when no worker pods exist. The HPA can therefore observe a persistent signal and trigger a scale‑up event.

  • Independence from pod existence: The metric source lives outside the pod set, so a zero replica count does not silence the metric.
  • Predictable scaling logic: The HPA can map a concrete value (e.g., 30 queued tasks) to a desired replica count.
  • Compatibility with durable queues: Work that can be buffered (e.g., message queues, task queues) tolerates the cold‑start latency introduced by scaling from zero.

Practical example: a queue consumer deployment named queue‑worker uses the external metric queue_consumer_lag exposed by the Prometheus Adapter. The metric is defined as the sum of pending tasks per queue name:

externalRules:
  - seriesQuery: '{__name__="queue_consumer_lag",name!=""}'
    metricsQuery: sum(<<.Series>>{<<.LabelMatchers>>}) by (name)
    resources:
      overrides:
        namespace:
          resource: namespace

After verifying the metric is reachable with:

kubectl get --raw '/apis/external.metrics.k8s.io/v1beta1/namespaces/default/queue_consumer_lag?labelSelector=name%3Dworker_tasks'

the HPA definition can be created:

apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: queue-worker
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: queue-worker
  minReplicas: 0
  maxReplicas: 10
  metrics:
  - type: External
    external:
      metric:
        name: queue_consumer_lag
        selector:
          matchLabels:
            name: worker_tasks
      target:
        type: Value
        value: "30"

When the queue is empty, the HPA reduces the deployment to zero replicas. As soon as tasks accumulate, the external metric remains available, the HPA computes the required replica count, and the controller schedules new pods. The ScaledToZero condition recorded by the HPA distinguishes an automatic scale‑down from a manual pause, ensuring that only workloads owned by the HPA are revived.

Step‑by‑Step: Configuring Prometheus Adapter for External Metrics

The Prometheus Adapter facilitates the translation of Prometheus time-series data into metrics compatible with the Kubernetes External Metrics API. By exposing metrics independently of Pod lifecycles, this adapter enables HorizontalPodAutoscaler (HPA) configurations to scale workloads—including those scaled down to zero—based on external signals like queue depth.

Configuring the adapter requires defining externalRules in your configuration manifest. These rules map Prometheus series to the API resources that the HPA controller consumes:

  • seriesQuery: A Prometheus query pattern used to discover available metrics. It should target the specific metric name while filtering for relevant labels to avoid excessive overhead.
  • metricsQuery: A template that the adapter executes to retrieve the current metric value. It utilizes <<.Series>> and <<.LabelMatchers>> placeholders to map the HPA's requested external metric label selectors to actual Prometheus queries.

A sample configuration snippet for a queue_consumer_lag metric follows:

externalRules:
- seriesQuery: '{__name__="queue_consumer_lag",name!=""}'
  metricsQuery: sum(<<.Series>>{<<.LabelMatchers>>}) by (name)
  resources:
    overrides:
      namespace:
        resource: namespace

Before deploying an HPA, you must verify that the metrics pipeline is correctly populating the External Metrics API. Use the kubectl get --raw command to query the API directly. This provides immediate feedback on whether the adapter successfully exposes the metric for a specific label selector:

kubectl get --raw '/apis/external.metrics.k8s.io/v1beta1/namespaces/default/queue_consumer_lag?labelSelector=name%3Dworker_tasks'

If this command fails or returns an empty result, the HPA will be unable to calculate replica counts, resulting in ScalingActive=False status. Only proceed to HPA configuration once this query returns a valid JSON payload containing the current metric value, confirming the observability bridge between Prometheus and the Kubernetes control plane is functional.

Defining the HPA: YAML Example and Behavior Details

Starting with Kubernetes v1.37, the HorizontalPodAutoscaler (HPA) natively supports scaling workloads to zero replicas using object or external metrics. This functionality is enabled by default, eliminating the need for custom add-ons. Because CPU and memory metrics are unavailable when no pods are running, this configuration requires metrics that exist independently of the workload, such as queue lengths.

apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: queue-worker
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: queue-worker
  minReplicas: 0
  maxReplicas: 10
  metrics:
  - type: External
    external:
      metric:
        name: queue_consumer_lag
        selector:
          matchLabels:
            name: worker_tasks
        target:
          type: Value
          value: "30"

The HPA utilizes a ScaledToZero status condition to differentiate between automatic scaling and manual intervention. When the HPA scales a target to zero, it sets ScaledToZero=True, indicating that the controller maintains ownership of the scaling logic and will continue to monitor metrics for scale-up triggers. If a user manually scales a deployment to zero, this condition is not set, and the HPA treats the workload as paused, refusing to perform further scaling actions.

Operational considerations include:

  • Stabilization Windows: The default downscale stabilization window remains five minutes. This prevents rapid toggling, though it can be adjusted via spec.behavior.scaleDown.
  • Metric Availability: The HPA cannot scale from zero if the metric source is unreachable. If the metrics adapter fails, the HPA will report ScalingActive=False.
  • Configuration Constraints: The API server rejects HPA manifests where minReplicas: 0 is combined solely with resource-based metrics (CPU/Memory).
  • Control Plane Versioning: During upgrades, ensure all control plane components support this feature before implementation, as older controllers may interpret minReplicas: 0 as a manual pause, leaving the workload inactive.

Upgrade, Compatibility, and Operational Considerations

The HPAScaleToZero feature gate is enabled by default on both the kube-apiserver and kube‑controller‑manager starting with Kubernetes v1.37. This default makes the minReplicas: 0 field in an HorizontalPodAutoscaler (HPA) a valid API contract, provided the HPA references an object or external metric.

Version‑skewed upgrade guidance

During a control‑plane upgrade where the API server and controller manager run different versions, follow these steps:

  • Confirm that the older component (typically the controller manager) has the HPAScaleToZero gate enabled; otherwise it will treat replicas: 0 as a manual pause.
  • Delay creation of HPAs with minReplicas: 0 until both components report the feature as active.
  • After the upgrade, verify the ScaledToZero condition on existing HPAs with kubectl describe hpa <name>.

Preparing to disable the gate or downgrade

Before turning off the feature gate or moving to a version that lacks the condition‑based implementation, perform the following:

  1. Update every affected HPA definition to set minReplicas to at least 1.
  2. Scale any workload currently at zero to a minimum of one replica (e.g., kubectl scale deployment <name> --replicas=1).
  3. Ensure each HPA still has a valid object or external metric; the API server will reject HPAs that rely solely on CPU or memory when minReplicas is zero.
  4. After the changes, you may safely disable the gate or roll back the cluster version.

Alpha‑to‑Beta evolution

The capability originated as an Alpha feature in Kubernetes v1.16. In v1.36 the ScaledToZero condition was added, allowing the controller to differentiate an automatic scale‑down from a manual pause. v1.37 graduated the feature to Beta and set the gate on by default, completing integration and end‑to‑end testing for external‑metric‑driven scaling to zero.

Practical example

apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: queue-worker
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: queue-worker
  minReplicas: 0
  maxReplicas: 10
  metrics:
  - type: External
    external:
      metric:
        name: queue_consumer_lag
        selector:
          matchLabels:
            name: worker_tasks
      target:
        type: Value
        value: "30"

Further learning resources

  • Official documentation: “Scaling to and from zero”.
  • KEP‑2021: “HPA supports scaling to and from zero pods for object and external metrics”.
  • Prometheus Adapter guide for exposing external metrics.
  • SIG Autoscaling community: Kubernetes Slack #sig-autoscaling channel.
APPWORKS ENGINEERING

Looking for Custom Software or AI Solutions?

Appworks Technologies designs, builds, and scales production enterprise platforms, microservices, and AI agent workflows tailored to your business goals.

Editorial Policy & Research Methodology

Our findings are based on rigorous internal research, verified industry benchmarks, and direct technical implementation experience from our enterprise client projects. All statistics and technical claims are reviewed by senior engineers before publication to ensure accuracy, transparency, and helpfulness for our readers.

Have an Idea? we offer services in Lucknow, Bangalore, Delhi NCR and other locations