Articles

Kubernetes v1.37: Scale Workloads to Zero with HorizontalPodAutoscaler

Kubernetes v1.37 introduces a Beta HorizontalPodAutoscaler that can automatically scale workloads down to zero using object or external metrics, saving resources while handling cold‑start considerations.

Written by:
APin

Senior Technology Analyst • Verified Expert

More from this author →
Kubernetes v1.37: Scale Workloads to Zero with HorizontalPodAutoscaler

Kubernetes v1.37 introduces a Beta HorizontalPodAutoscaler that can automatically scale workloads down to zero using object or external metrics, saving resources while handling cold‑start considerations.

Beta HPA Scaling to Zero in Kubernetes v1.37

The Beta HorizontalPodAutoscaler (HPA) in Kubernetes v1.37 adds native support for scaling a workload down to zero replicas and back up again without requiring an external add‑on or an Alpha feature gate. This capability is enabled by default on both the kube‑apiserver and the kube‑controller‑manager, making the API field minReplicas: 0 a first‑class option.

Why object or external metrics are required

Traditional HPA scaling relies on CPU or memory usage, which are only emitted while Pods are running. When the replica count reaches zero there are no Pods to produce those metrics, so the controller loses its signal. Object and external metrics, such as a queue length, exist independently of the workers and can be read continuously. The HPA therefore continues to evaluate the metric and can trigger a scale‑up even when no Pods are present.

Practical configuration example

Assume a Prometheus metric queue_consumer_lag that reports the number of pending tasks. After installing a Prometheus Adapter that exposes this series via the External Metrics API, the following steps create a zero‑to‑ten scaling policy for a Deployment named queue-worker:

apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: queue-worker
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: queue-worker
  minReplicas: 0
  maxReplicas: 10
  metrics:
  - type: External
    external:
      metric:
        name: queue_consumer_lag
        selector:
          matchLabels:
            name: worker_tasks
      target:
        type: Value
        value: "30"

When the queue is empty, the HPA reduces the Deployment to zero. As soon as the metric reports ≥ 30 pending tasks, the controller calculates the required replica count (capped at ten) and schedules new Pods.

Operational considerations

  • Cold‑start latency: Scaling from zero incurs the time to observe the metric, schedule a Pod, and start the application. Workloads that can tolerate this delay—e.g., batch processors or queue consumers—are good candidates.
  • Stabilization window: The default downscale stabilization period is five minutes, preventing brief metric drops from instantly removing all workers. Adjust spec.behavior.scaleDown if a different window is required.
  • Paused vs. scaled‑to‑zero: The controller records a ScaledToZero=True condition when it performs the scale‑down. If an operator manually sets replicas to zero, the condition remains false, and the HPA will not auto‑scale back up.

Before upgrading a control‑plane with mixed component versions, ensure both the API server and controller manager have the feature enabled; otherwise a zero replica count may be interpreted as a manual pause. When downgrading or disabling the feature gate, change any HPA with minReplicas: 0 to a minimum of one and verify that at least one object or external metric is defined, as the API server will reject HPAs that rely solely on resource metrics.

Choosing the Right Metric: Object and External Metrics

HorizontalPodAutoscaler (HPA) implementations that rely on resource metrics such as CPU or memory are inherently limited when a workload is scaled to zero. Both metrics are emitted only by running Pods; once the replica count reaches zero there are no Pods to produce the telemetry, so the HPA loses its signal source. Without a metric, the controller cannot determine whether new work has arrived, and therefore cannot trigger a scale‑up.

Object and external metrics solve this problem because they are decoupled from the lifecycle of the Pods that consume them. An object metric (e.g., the length of a Kafka topic) or an external metric (e.g., a Prometheus series named queue_consumer_lag) exists independently of the worker processes. The HPA can continue to poll these values even when replicas: 0, allowing it to detect demand and create new Pods on demand.

  • Metric availability at zero: Object/external metrics are produced by the system that generates work (a queue, a database, a monitoring exporter), not by the consumer Pods.
  • Deterministic scaling logic: The HPA can map a concrete value (e.g., 30 queued tasks) to a replica count, as shown in the example where one replica is added per 30 tasks.
  • Clear ownership of zero state: Kubernetes v1.37 introduces the ScaledToZero condition, which distinguishes an automatic scale‑down from a manual pause, ensuring the HPA resumes when the external metric signals work.

Practical example: a queue‑consumer deployment uses the Prometheus Adapter to expose queue_consumer_lag via the External Metrics API. The HPA definition includes:

apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: queue-worker
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: queue-worker
  minReplicas: 0
  maxReplicas: 10
  metrics:
  - type: External
    external:
      metric:
        name: queue_consumer_lag
        selector:
          matchLabels:
            name: worker_tasks
      target:
        type: Value
        value: "30"

When the queue is empty, the external metric reports zero, and the HPA reduces the deployment to zero replicas, releasing expensive resources such as dedicated CPUs or GPUs. As soon as tasks accumulate, the metric remains available, the HPA computes the required replica count, and the workload is revived without manual intervention.

Setting Up External Metrics with Prometheus Adapter

External metrics allow the Horizontal Pod Autoscaler (HPA) to make scaling decisions based on data that exists outside of a Pod, such as queue length or custom business indicators. Kubernetes exposes these metrics through the External Metrics API, but a metrics provider must translate a monitoring system’s data model into the API format. The Prometheus Adapter is the most common provider for Prometheus‑based environments.

Adapter configuration

Define an externalRules entry that maps a Prometheus series to an external metric name. The rule consists of a seriesQuery to select the raw series, a metricsQuery that aggregates the series, and optional resources overrides to bind the metric to a Kubernetes resource type.

apiVersion: v1
kind: ConfigMap
metadata:
  name: prometheus-adapter-config
data:
  config.yaml: |
    externalRules:
    - seriesQuery: '{__name__="queue_consumer_lag",name!=""}'
      metricsQuery: sum(<<.Series>>{<<.LabelMatchers>>}) by (name)
      resources:
        overrides:
          namespace:
            resource: namespace

Adjust the seriesQuery to match the exact label set used by your Prometheus scrape target. The metricsQuery can be any valid PromQL expression; the example aggregates lag per queue name.

Verify metric availability

Before creating an HPA, confirm that the API server can retrieve the metric:

kubectl get --raw '/apis/external.metrics.k8s.io/v1beta1/namespaces/default/queue_consumer_lag?labelSelector=name%3Dworker_tasks'

The response should contain the current value for the worker_tasks label. If the call fails, troubleshoot the Prometheus scrape, the adapter’s discovery rules, or network policies that may block the request.

Installation considerations

  • Version compatibility: The adapter must be deployed on a cluster where the HPAScaleToZero feature gate is enabled (default in Kubernetes v1.37). Both the API server and controller manager need the gate active for minReplicas: 0 to work.
  • Discovery rules: The adapter’s rules section may need additional resource overrides if you expose metrics for custom resources or namespaces.
  • Security posture: Follow your organization’s compliance framework (e.g., ISO 27001 or NIST) by restricting the adapter’s ServiceAccount to read‑only access on the Prometheus endpoint and limiting its network egress.
  • Observability: Enable the adapter’s own metrics (exposed on /metrics) and forward them to your monitoring stack to detect metric‑lookup failures that would cause the HPA to report ScalingActive=False.

Once the adapter is verified, you can reference the external metric in an HPA definition, allowing workloads such as queue consumers to scale to zero when the metric reports no pending work and automatically scale back up when tasks appear.

Creating an HPA that Scales to Zero

Scaling a queue‑consumer workload to zero requires an external metric that remains observable when no Pods are running. The queue_consumer_lag metric collected by Prometheus satisfies this requirement because it reflects the length of a durable queue independent of the workers.

Step‑by‑step HPA definition

  1. Verify metric exposure through the External Metrics API:

    kubectl get --raw '/apis/external.metrics.k8s.io/v1beta1/namespaces/default/queue_consumer_lag?labelSelector=name%3Dworker_tasks'

    The command must return a numeric value; otherwise fix the Prometheus Adapter configuration.

  2. Create a Deployment that can run with at least one replica. This initial replica allows the HPA controller to own the scaling lifecycle.

    apiVersion: apps/v1
    kind: Deployment
    metadata:
      name: queue-worker
    spec:
      replicas: 1
      selector:
        matchLabels:
          app: queue-worker
      template:
        metadata:
          labels:
            app: queue-worker
        spec:
          containers:
          - name: worker
            image: your-registry/queue-worker:latest
            resources:
              requests:
                cpu: "250m"
                memory: "256Mi"
  3. Apply the HPA that permits scaling from zero to ten replicas, using one replica per 30 queued tasks:

    apiVersion: autoscaling/v2
    kind: HorizontalPodAutoscaler
    metadata:
      name: queue-worker
      annotations:
        kubernetes.io/description: "Scales queue-worker based on the number of queued tasks"
    spec:
      scaleTargetRef:
        apiVersion: apps/v1
        kind: Deployment
        name: queue-worker
      minReplicas: 0
      maxReplicas: 10
      metrics:
      - type: External
        external:
          metric:
            name: queue_consumer_lag
            selector:
              matchLabels:
                name: worker_tasks
          target:
            type: Value
            value: "30"

Stabilization window and scaling behavior

  • The default downscale stabilization window is five minutes. During this period a brief dip in queue_consumer_lag will not immediately reduce the replica count to zero, preventing unnecessary cold‑starts.
  • If a different latency profile is needed, configure spec.behavior.scaleDown.stabilizationWindowSeconds to a custom value.
  • When the HPA reduces the Deployment to zero, it sets the ScaledToZero=True condition. This condition tells the controller that the zero state was produced automatically, so it continues to poll the external metric.
  • When traffic arrives and the metric exceeds the target, the controller clears the condition (ScaledToZero=False) and creates the required Pods.
  • Manual scaling to zero (e.g., kubectl scale deployment queue-worker --replicas=0) does not set the condition, leaving the workload paused until an operator intervenes.

Before upgrading a control‑plane version, ensure that both the API server and controller manager have the HPAScaleToZero feature gate enabled; otherwise a replica count of zero will be interpreted as a manual pause and the workload will remain idle.

Operational Tips, Upgrade Path, and Troubleshooting

The HPAScaleToZero feature gate is enabled by default in Kubernetes v1.37 on both the kube‑apiserver and the kube‑controller‑manager. When the gate is active, the API server accepts minReplicas: 0 and the controller manager adds a ScaledToZero condition to the HPA status. This condition distinguishes an automatic scale‑down (the controller set the replica count to zero) from a manual pause (an operator set the replica count to zero). The controller only continues to evaluate object or external metrics while the condition is true.

Version‑skewed upgrade guidance

  • During a control‑plane upgrade, ensure that both the API server and the controller manager are running a version that has the feature gate enabled before creating HPAs with minReplicas: 0.
  • If a controller manager without the gate is present, a replica count of zero is interpreted as a manual pause, leaving the workload stopped.
  • Before downgrading to a version that lacks the condition‑based implementation, modify existing HPAs to minReplicas: 1 and scale any zero‑replica workloads to at least one pod.

Distinguishing automatic zero from manual pause

After the HPA scales a Deployment to zero, it records ScaledToZero=True. Operators can verify this with:

kubectl describe hpa queue-worker

If the condition is absent, the zero state is considered paused. Resuming the workload requires an explicit kubectl scale --replicas=1 deployment/queue-worker or fixing the metric source.

Handling missing external metrics

When the configured external metric cannot be retrieved, the HPA reports ScalingActive=False with a reason such as FailedGetExternalMetric. The recommended remediation steps are:

  1. Validate the metrics pipeline (e.g., kubectl get --raw '/apis/external.metrics.k8s.io/v1beta1/...').
  2. Restore the metric source or adjust the HPA to use a fallback metric.
  3. Manually scale the workload to restore capacity while the metric issue is resolved.

Best‑practice recommendations

  • Always pair a zero‑scale HPA with an object or external metric (queue length, Prometheus queue_consumer_lag, etc.); pure CPU/memory metrics cannot drive a scale‑to‑zero.
  • Configure a downscale stabilization window appropriate to your workload (default five minutes) via spec.behavior.scaleDown to avoid rapid oscillations.
  • Document the expected pause behavior in runbooks and include checks for the ScaledToZero condition before performing manual scaling actions.
  • During upgrades, verify feature‑gate status on both control‑plane components before enabling minReplicas: 0 in new HPAs.
APPWORKS ENGINEERING

Looking for Custom Software or AI Solutions?

Appworks Technologies designs, builds, and scales production enterprise platforms, microservices, and AI agent workflows tailored to your business goals.

Editorial Policy & Research Methodology

Our findings are based on rigorous internal research, verified industry benchmarks, and direct technical implementation experience from our enterprise client projects. All statistics and technical claims are reviewed by senior engineers before publication to ensure accuracy, transparency, and helpfulness for our readers.

Have an Idea? we offer services in Lucknow, Bangalore, Delhi NCR and other locations