Articles

Kubernetes v1.37 Introduces Scheduler Preemption for In-Place Pod Resize (Alpha)

Kubernetes v1.37 adds an alpha feature that lets the scheduler preempt lower‑priority pods to satisfy in‑place pod resize requests that were previously deferred. This bridges the resource scheduling gap introduced by the GA in‑place resize feature in v1.35.

Written by:
APin

Senior Technology Analyst • Verified Expert

More from this author →
Kubernetes v1.37 Introduces Scheduler Preemption for In-Place Pod Resize (Alpha)

Kubernetes v1.37 adds an alpha feature that lets the scheduler preempt lower‑priority pods to satisfy in‑place pod resize requests that were previously deferred. This bridges the resource scheduling gap introduced by the GA in‑place resize feature in v1.35.

Background: Evolution of In-Place Pod Resize

Prior to the General Availability (GA) of the core in‑place pod resize feature in Kubernetes v1.35, a pod’s CPU and memory limits were immutable after the initial scheduling decision. The scheduler placed the pod based on the requested resources, and any subsequent change required deleting the pod and creating a new one, which caused a brief service interruption.

With GA in‑place resize, the kubelet can apply a new resources.requests or resources.limits to a running container without terminating the pod. The workflow is:

  1. A controller (e.g., the Vertical Pod Autoscaler) updates the pod spec with higher CPU or memory values.
  2. The kubelet checks the node’s allocatable capacity. If sufficient headroom exists, it updates the cgroup limits and reports the new values in status.containerStatuses[].resources.
  3. If the node lacks free capacity, the kubelet marks the resize as Deferred rather than rejecting it outright. The pod remains running, awaiting resource availability.

This mechanism preserves the “no‑restart” guarantee for workloads that need to react to sudden load spikes, such as an in‑memory cache or a latency‑sensitive API server.

  • Dynamic scaling: Applications can increase memory to avoid OOM termination without downtime.
  • Operational simplicity: Operators no longer need to manually evict lower‑priority pods or trigger a full node replacement to accommodate a resize.
  • Resource efficiency: Nodes can be packed with best‑effort workloads, knowing that critical pods can request additional resources on‑the‑fly.

Because the resize request is deferred when the node is fully utilized, administrators still have options to unblock it:

  • Manually evict low‑priority pods from the node.
  • Allow the cluster autoscaler to provision a larger node and migrate the pod.
  • Enable scheduler‑driven preemption (introduced as an alpha feature in v1.37) to automatically evict lower‑priority pods on the same node.

Example of a manual resize request using kubectl:

kubectl patch pod my-app -p '{"spec":{"containers":[{"name":"app","resources":{"requests":{"memory":"2Gi"}}}]}}' --type=merge

After the patch, the kubelet either applies the new limit immediately or sets status.containerStatuses[].resizeStatus = "Deferred". Monitoring this status with kubectl get pod my-app -o yaml helps operators understand whether additional capacity must be freed before the resize can be completed.

The Deferred Resize Challenge

When a controller such as the Vertical Pod Autoscaler updates the resources.requests of a running container, the Kubelet first checks the node’s allocatable headroom. If the node cannot satisfy the increase, the Kubelet does not reject the request outright. Instead it marks the container’s resizeStatus as Deferred and leaves the Pod in a waiting state until enough resources become free.

This behavior differs from an Infeasible resize request. An infeasible request is rejected immediately because it violates hard limits such as the physical machine capacity, namespace limit ranges, or admission‑quota constraints. A Deferred status, by contrast, signals that the request is valid but temporarily un‑actuable because the node is fully utilized.

Before scheduler‑driven preemption was introduced, Deferred resizes created several operational pain points:

  • Manual eviction cycles – administrators had to identify lower‑priority Pods on the same node and evict them manually to free headroom, a process prone to human error and service disruption.
  • Reliance on cluster autoscaler – the only automated path was to add a larger node and reschedule the Pod, which contradicted the “no restart” promise of in‑place scaling and introduced latency.
  • Potential permanent blockage – if a node remained saturated, a critical workload (e.g., an in‑memory database facing an imminent OOM) could stay Deferred indefinitely, risking application failure.
  • Resource fragmentation – operators often left unused buffer capacity on nodes to avoid Deferred states, reducing overall cluster utilization and increasing cost.

A practical example illustrates the issue: a real‑time web server running on a node with 4 GiB of allocatable memory receives a VPA recommendation to increase its limit to 5 GiB. The node already hosts several batch jobs consuming the remaining 1 GiB. The Kubelet marks the resize as Deferred. Without preemption, the only remediation is to manually kill one or more batch jobs or wait for the autoscaler to provision a new node, both of which interrupt the intended seamless scaling.

These challenges highlighted the need for a centralized preemption mechanism that can automatically evict lower‑priority Pods on the same node, allowing Deferred resizes to complete without manual intervention or unnecessary node expansion.

Introducing Scheduler Preemption for In-Place Resize (Alpha)

The InPlacePodVerticalScalingSchedulerPreemption feature gate activates a new preemption path in the kube‑scheduler that is dedicated to Pods whose in‑place resize request is marked Deferred. When a running Pod’s resizeStatus is set to Deferred, the scheduler treats the Pod as still “active” in the scheduling cycle instead of assuming it is fully placed.

Key mechanics:

  • Centralized tracking: The scheduler continuously watches the cluster for Pods with the Deferred condition. These Pods remain in the active scheduling queue until the Kubelet reports a successful resize.
  • Node‑local preemption scope: Preemption is limited to the node where the deferred Pod is already running. The scheduler identifies lower‑priority “victim” Pods on that same node, evicts them gracefully, and thereby releases the required CPU or memory.
  • Resource reservation safety: The scheduler assumes the resources requested for the resize are already consumed. This prevents double‑allocation races and ensures the Kubelet can apply the resize immediately after eviction completes.
  • Separation of concerns: The Kubelet’s critical‑Pod admission handler no longer performs local preemption for in‑place resizes. All decisions are centralized in the scheduler, which respects global priorities, PodDisruptionBudgets, and termination policies.

Example scenario

# A high‑priority web server requests +2Gi memory
kubectl patch pod web-frontend -p '{"spec":{"containers":[{"name":"web","resources":{"requests":{"memory":"4Gi"}}}]}}' --type=merge

# The node is fully utilized; the Kubelet marks the request Deferred
# Scheduler sees the Deferred pod, selects a low‑priority batch job on the same node,
# evicts it, and the memory becomes available for the web server’s resize.

Administrators can disable this behavior per node using spec.podPreemptionPolicy with the disableResizePreemption flag, allowing custom autoscaling logic to take precedence.

To enable the feature, the cluster must run Kubernetes v1.37 or later and have the feature gate turned on for kube-apiserver, kube-scheduler, and the Kubelet. Once active, the scheduler automatically resolves Deferred resize requests without manual eviction or disruptive node scaling, preserving the “no‑restart” guarantee of in‑place vertical scaling while maintaining high node utilization.

Architectural Mechanics and Safety Guarantees

Kubernetes v1.37 introduces a dedicated preemption path for in‑place pod resizing. The core of this mechanism is a centralized scheduler tracking loop that watches for Pods whose status.containerStatuses[].resizeStatus is set to Deferred. Normally a Pod with spec.nodeName is excluded from the active scheduling queue, but the feature gate InPlacePodVerticalScalingSchedulerPreemption forces the scheduler to keep these Pods in the queue until the Kubelet reports a successful resize.

Single‑node preemption boundary limits the eviction scope to the node where the deferred Pod resides. The scheduler enumerates lower‑priority “victim” Pods on that host, respects their PodDisruptionBudget and graceful termination policies, and issues eviction requests. If the node cannot free enough resources even after evicting all eligible Pods, the resize remains Deferred and no cross‑node migration is attempted.

Resource reservation safety is achieved by treating the requested resize amount as already allocated in the scheduler’s internal snapshot. This prevents race conditions where two resize operations could simultaneously claim the same headroom, leading to double‑allocation.

Separation of concerns with the Kubelet’s critical‑Pod admission handler ensures that local eviction logic is not duplicated for resize events. The Kubelet still uses its critical‑Pod handler for new pod admission, but for in‑place resizing it defers entirely to the scheduler. This centralization guarantees that global priority ordering, PDB compliance, and cluster‑wide policies are applied consistently.

Competing resize requests are handled dynamically. When a higher‑priority resize arrives on the same node during an ongoing preemption cycle, the Kubelet reports the new request as Deferred. The scheduler re‑evaluates the node’s snapshot, selects additional victims if needed, and triggers a second eviction round. This ensures that the most critical workloads always obtain the required headroom.

  • Enable the feature gate on kube-apiserver, kube-scheduler, and kubelet.
  • Observe Deferred pods with kubectl get pod -o jsonpath='{.status.containerStatuses[*].resizeStatus}'.
  • Configure node‑level opt‑out via spec.podPreemptionPolicy.disableResizePreemption when preemption is undesirable.

By confining preemption to the originating node, reserving resources ahead of actuation, and delegating all decision‑making to the scheduler, Kubernetes v1.37 provides a deterministic and safe pathway for critical workloads to scale without restarts, while preserving high bin‑packing efficiency for lower‑priority jobs.

Getting Started: Enabling and Testing the Feature

Before enabling the scheduler preemption for in‑place pod resize, understand that the feature is gated by InPlacePodVerticalScalingSchedulerPreemption and is available only in Kubernetes v1.37 or later. The scheduler watches Pods whose status.containerStatuses[].resizeStatus is set to Deferred and, if necessary, evicts lower‑priority Pods on the same node to free capacity for the pending resize.

Enable the feature gate on control‑plane components

  • Edit the static pod manifests (usually under /etc/kubernetes/manifests) for kube-apiserver, kube-scheduler, and kube-controller-manager.
  • Add the flag --feature-gates=InPlacePodVerticalScalingSchedulerPreemption=true to the command array of each manifest.
  • Restart the static pods (or simply let the kubelet reload the manifests) so the new flag takes effect.

Enable the feature gate on the kubelet

  • If you use a KubeletConfiguration file, set featureGates.InPlacePodVerticalScalingSchedulerPreemption: true and restart the kubelet service.
  • When using command‑line flags, add --feature-gates=InPlacePodVerticalScalingSchedulerPreemption=true to the kubelet startup options.

Configure node‑level preemption policies (optional)

To disable resize preemption on specific nodes, patch the node object with the new spec.podPreemptionPolicy field:

apiVersion: v1
kind: Node
metadata:
  name: batch-workload-node
spec:
  podPreemptionPolicy:
    disableResizePreemption:
    - "cluster-autoscaler.kubernetes.io/disable-preemption"
    - "operator.example.com/policy-override"

This tells the scheduler not to preempt lower‑priority Pods for in‑place resizes on that node, allowing controllers to manage capacity themselves.

Mini‑tutorial on a kind cluster

  1. Create a kind-config.yaml that enables the feature gate:
    # kind-config.yaml
    kind: Cluster
    apiVersion: kind.x-k8s.io/v1alpha4
    featureGates:
      InPlacePodVerticalScalingSchedulerPreemption: true
    
  2. Provision the cluster with a v1.37 node image:
    kind create cluster --config kind-config.yaml --image kindest/node:v1.37.0
    
  3. Deploy a low‑priority workload that consumes most of the node’s CPU:
    kubectl run low‑priority --image=busybox --restart=Never --requests=cpu=900m --priorityClassName=low-priority
    
  4. Deploy a higher‑priority Pod that will request an in‑place resize:
    kubectl run high‑priority --image=busybox --restart=Never --requests=cpu=200m --priorityClassName=high-priority
    # Then patch it to request more CPU
    kubectl patch pod high‑priority -p '{"spec":{"containers":[{"name":"busybox","resources":{"requests":{"cpu":"1200m"}}}]}}' --type=merge
    
  5. Observe the resize status:
    kubectl get pod high‑priority -o jsonpath='{.status.containerStatuses[0].resizeStatus}'
    
    If the status changes from Deferred to a successful resize, the scheduler has preempted the low‑priority Pod on the same node.

These steps demonstrate how to activate and verify scheduler‑driven preemption for in‑place pod resize, enabling higher‑priority workloads to obtain needed resources without manual eviction.

APPWORKS ENGINEERING

Looking for Custom Software or AI Solutions?

Appworks Technologies designs, builds, and scales production enterprise platforms, microservices, and AI agent workflows tailored to your business goals.

Editorial Policy & Research Methodology

Our findings are based on rigorous internal research, verified industry benchmarks, and direct technical implementation experience from our enterprise client projects. All statistics and technical claims are reviewed by senior engineers before publication to ensure accuracy, transparency, and helpfulness for our readers.

Have an Idea? we offer services in Lucknow, Bangalore, Delhi NCR and other locations