Articles

Kubernetes v1.37: New Node Lifecycle Conditions Explained

Kubernetes v1.37 adds five well‑known Node lifecycle conditions—DrainInProgress, Drained, MaintenancePlanned, MaintenanceInProgress, and GracefulNodeShutdownInProgress—giving operators a shared, Kubernetes‑owned signal for drains, maintenance windows, and graceful shutdowns.

Written by:
APin

Senior Technology Analyst • Verified Expert

More from this author →
Kubernetes v1.37: New Node Lifecycle Conditions Explained

Kubernetes v1.37 adds five well‑known Node lifecycle conditions—DrainInProgress, Drained, MaintenancePlanned, MaintenanceInProgress, and GracefulNodeShutdownInProgress—giving operators a shared, Kubernetes‑owned signal for drains, maintenance windows, and graceful shutdowns.

Overview of Node Lifecycle Conditions

Node‑level operations such as draining, hardware upgrades, or graceful shutdown affect many cluster components. The kubelet, scheduler, autoscalers, storage operators, and custom maintenance controllers each need to know whether a node is simply “NotReady” or is deliberately being taken out of service. Without a common, Kubernetes‑owned signal, each component infers state from disparate sources—readiness, taints, pod termination, or provider‑specific annotations—leading to contradictory actions (e.g., a DaemonSet controller recreating a pod that the kubelet is shutting down).

Introducing a shared signal solves this coordination problem. The NodeLifecycleConditions feature gate reserves five well‑known condition types that any authorized controller can set on node.status.conditions. The conditions are machine‑readable, have a stable reason, and can carry a human‑readable message for dashboards or alerts.

  • DrainInProgress – Indicates that the node is actively being drained according to the administrator’s criteria (e.g., kubectl drain with custom eviction thresholds). Controllers can pause new pod placements while the condition is True.
  • Drained – Signals that the node has satisfied the chosen drain criteria. At this point, workload controllers may safely stop scheduling new pods, and external automation can proceed to the next maintenance step.
  • MaintenancePlanned – Communicates a future maintenance window (hardware replacement, OS patch, etc.). An automation system might publish this condition weeks in advance, allowing capacity planners to adjust replica counts or trigger pre‑maintenance alerts.
  • MaintenanceInProgress – Marks that the node is currently undergoing maintenance, whether it requires a drain or not (e.g., firmware upgrade or debugging session). Components that respect this signal can avoid triggering unnecessary evictions.
  • GracefulNodeShutdownInProgress – Declares that the node is executing a graceful shutdown sequence, enabling workload controllers to wait for pod termination rather than treating the node as a failure.

Practical usage example:

kubectl patch node worker-01 -p '{
  "status": {
    "conditions": [
      {"type":"MaintenancePlanned","status":"True","reason":"HWWindow","message":"Scheduled hardware replacement"}
    ]
  }
}' --type=merge

Later, when the maintenance window opens, the same controller updates the condition to MaintenanceInProgress, then to DrainInProgress if a drain is required, and finally to Drained once the criteria are met. By publishing these states in a single, canonical location, the cluster gains operational clarity and reduces the risk of conflicting decisions across its diverse controllers.

Detailed Breakdown of Each Condition

The Kubernetes v1.37 Node lifecycle model defines five well‑known NodeConditionType values that convey a node’s operational state in a machine‑readable way. Each condition follows the standard condition schema: a status field (True, False, Unknown), a stable reason string, and an optional human‑readable message. Controllers or administrators set these fields to make drain, maintenance, and shutdown activities visible to other components, dashboards, and alerting pipelines.

DrainInProgress
  • Status: True while the node is actively being drained according to the administrator‑defined criteria; False or omitted when no drain is occurring; Unknown if the kubelet cannot determine the drain state.
  • Reason: Typically a short identifier such as PodEviction or NodeCordon that indicates why the drain was initiated.
  • Message: Free‑form text, e.g., “Evicting 12 Pods to prepare for hardware upgrade”.
Drained
  • Status: True once the node satisfies the selected drain criteria (e.g., no non‑daemonset pods remain); False or omitted otherwise; Unknown when the condition cannot be evaluated.
  • Reason: DrainComplete or a custom token that signals successful completion.
  • Message: “All evictable workloads have terminated; node ready for maintenance”.
MaintenancePlanned
  • Status: True when a future maintenance window is scheduled; False when no such window exists; Unknown if the schedule source is unreachable.
  • Reason: MaintenanceWindow (as shown in the blog example).
  • Message: “Hardware maintenance is scheduled for this node at 2026‑12‑09T12:00:00Z”.
MaintenanceInProgress
  • Status: True while the node is undergoing any maintenance activity (hardware replacement, software rollout, debugging); False when maintenance is not active; Unknown if the controller cannot confirm the state.
  • Reason: A stable token such as HardwareSwap, SoftwareUpgrade, or DebugSession.
  • Message: “Applying kernel live‑patch; drain not required”.
GracefulNodeShutdownInProgress
  • Status: True when the node is executing a graceful shutdown sequence; False when the node is running normally; Unknown if the shutdown controller is unreachable.
  • Reason: GracefulShutdown or a provider‑specific identifier.
  • Message: “Node is terminating pods with pre‑stop hooks before power‑off”.

In practice, an automation script might set MaintenancePlanned with status: "True" and reason: "MaintenanceWindow" weeks before work begins, then flip MaintenanceInProgress to True at the start of the window, and finally clear both conditions after the Drained condition reaches True. Using these shared signals reduces the need for ad‑hoc annotations and enables consistent decision‑making across the scheduler, autoscalers, and storage operators.

How to Publish and Use Conditions Today

Kubernetes v1.37 defines five well‑known NodeCondition types (e.g., DrainInProgress, MaintenancePlanned) that an administrator or an authorized controller can publish directly on a Node object. The conditions are part of the Node’s .status.conditions array and use the standard status field with values True, False, or Unknown. A stable reason string provides a machine‑readable cause, while message supplies free‑form human context.

Operational steps to set a condition

  1. Determine ownership of the condition (e.g., a maintenance automation controller owns MaintenancePlanned).
  2. Construct a patch that adds or updates the condition with status: "True" and a concise, immutable reason (e.g., ScheduledWindow).
  3. Apply the patch with kubectl patch node <node-name> --type=merge -p … or via the Kubernetes API from a controller.
  4. When the lifecycle state ends, either set status: "False" or remove the condition entirely to avoid stale signals.

Example patch that marks a node as draining:

kubectl patch node worker-01 \
  --type=merge -p '{
    "status": {
      "conditions": [
        {
          "type": "DrainInProgress",
          "status": "True",
          "reason": "AdminDrain",
          "message": "Drain started by ops team"
        }
      ]
    }
  }'

To clear the condition after the drain criteria are met:

kubectl patch node worker-01 \
  --type=merge -p '{
    "status": {
      "conditions": [
        {
          "type": "DrainInProgress",
          "status": "False",
          "reason": "Completed",
          "message": "All pods evicted"
        }
      ]
    }
  }'

Recommended use of True/False and reason strings

  • Set status: "True" only while the lifecycle event is actively observed.
  • Use a short, immutable reason (e.g., ScheduledWindow, AdminDrain) so automation can reliably filter on it.
  • Provide a human‑readable message for dashboards or alerts.
  • When the event ends, either set status: "False" with a new reason or delete the condition to prevent false positives.

Integration with existing tools

  • kubectl cordon and kubectl drain continue to control scheduling and eviction; they do not set the new conditions automatically.
  • Apply taints (e.g., node.kubernetes.io/unschedulable) to block new pods while a condition such as DrainInProgress is True.
  • Automation that runs kubectl drain should publish DrainInProgress before eviction and Drained after the configured criteria are satisfied.
  • Maintenance workflows that use kubectl cordon can also set MaintenanceInProgress to give a shared, observable signal to other controllers (e.g., DaemonSet, Job, storage operators).

By publishing conditions independently of the scheduling/eviction mechanisms, administrators obtain a single, authoritative source of node lifecycle state that can be consumed by dashboards, alerting pipelines, and future core controllers without altering existing operational practices.

Impact on Cluster Components and Operational Clarity

The kubelet, scheduler, DaemonSet controller, Job controller, storage operators, and other core components each infer a node’s state from disparate signals such as Ready status, taints, or pod termination events. Without a shared, authoritative indicator, these components can make contradictory decisions during maintenance or graceful shutdown.

Node lifecycle conditions introduced in Kubernetes v1.37—DrainInProgress, Drained, MaintenancePlanned, MaintenanceInProgress, and GracefulNodeShutdownInProgress—provide a single, Kubernetes‑owned place on the Node.status.conditions array to publish the exact lifecycle phase. When a controller or automation sets a condition to True, the reason and message fields give a stable, machine‑readable context that other components can consume.

Practical impact on key controllers:

  • Kubelet: Detects GracefulNodeShutdownInProgress and refrains from restarting pods that are being terminated, avoiding a race where a newly started pod is immediately evicted.
  • Scheduler: Checks MaintenanceInProgress or DrainInProgress before considering a node for new pod placement, preventing scheduling onto a node that is about to lose capacity.
  • DaemonSet controller: Uses MaintenanceInProgress to exclude a node from rollout calculations, ensuring that a node taken offline does not consume the availability budget and block progress on healthy nodes.
  • Job controller: Observes DrainInProgress and can mark unfinished pods as failed or reschedule them, rather than waiting indefinitely for a pod that will never reach a terminal phase.
  • Storage operators: React to MaintenancePlanned by pre‑emptively relocating volumes or adjusting replication factors before a drain begins, reducing I/O disruption.

By consolidating lifecycle intent into these conditions, administrators eliminate the need for each component to reconstruct state from indirect cues. The result is operational clarity: alerts, dashboards, and automation can all read the same signal, and ownership of each condition can be assigned to avoid conflicting writes. This shared context is the foundation for future enhancements—such as condition‑aware rollout ordering or automated lock acquisition—while preserving existing mechanisms like kubectl cordon, taints, and eviction policies.

Future Roadmap and Community Involvement

The Alpha NodeLifecycleConditions feature gate was introduced in Kubernetes v1.37 as a placeholder for a shared, Kubernetes‑owned signal describing node‑level activities such as draining, maintenance, or graceful shutdown. In the initial release the gate is a no‑op: it does not restrict who can set the conditions, and no core controller reads them. This design allows administrators or authorized automation to publish conditions today while preserving the ability to opt‑in to future controller behavior without a disruptive upgrade.

Practical example – publishing a maintenance condition

apiVersion: v1
kind: Node
metadata:
  name: worker-01
status:
  conditions:
  - type: MaintenancePlanned
    status: "True"
    reason: MaintenanceWindow
    message: "Hardware maintenance scheduled for this node"
    lastTransitionTime: "2026-12-09T12:00:00Z"

Administrators should set status: "True" while the lifecycle state is active and clear or set to "False" once it ends. Using a stable reason and concise message enables both humans and automation to interpret the signal reliably.

Upcoming core controller consumption

  • DaemonSet controller: will eventually ignore Pods on nodes marked MaintenanceInProgress for rollout budgeting, reducing false‑positive unavailability.
  • Job controller: will treat GracefulNodeShutdownInProgress as a normal termination path, preventing indefinite waits.
  • Scheduler and autoscalers: will incorporate DrainInProgress and Drained to avoid scheduling new workloads onto nodes that are being evacuated.

Node Lifecycle Working Group and KEP‑5683

The Node Lifecycle Working Group, operating under SIG Node and SIG Apps, steers the evolution of these conditions. KEP‑5683 documents the design, well‑known condition types, and the roadmap for controller integration. The group regularly publishes design proposals, gathers use‑case feedback, and coordinates with related projects such as cluster‑autoscaler, storage operators, and fleet‑management tools.

How engineers can contribute

  • Review KEP‑5683 on the Kubernetes enhancement tracker and comment on open design questions.
  • Submit a pull request that adds a controller‑side consumer for a specific condition (e.g., extending the DaemonSet controller to respect MaintenanceInProgress).
  • Share operational patterns in the SIG Node mailing list or the SIG Apps forum to help define ownership and locking semantics.
  • Participate in the quarterly Node Lifecycle Working Group meetings to align on cross‑component expectations.

By publishing conditions today and collaborating on the next set of controller enhancements, the community can gradually replace ad‑hoc signals (taints, annotations, readiness) with a single, authoritative source of node lifecycle state.

APPWORKS ENGINEERING

Looking for Custom Software or AI Solutions?

Appworks Technologies designs, builds, and scales production enterprise platforms, microservices, and AI agent workflows tailored to your business goals.

Editorial Policy & Research Methodology

Our findings are based on rigorous internal research, verified industry benchmarks, and direct technical implementation experience from our enterprise client projects. All statistics and technical claims are reviewed by senior engineers before publication to ensure accuracy, transparency, and helpfulness for our readers.

Have an Idea? we offer services in Lucknow, Bangalore, Delhi NCR and other locations