The new release lets GPU‑intensive workloads shrink to zero pods when not in use, instantly releasing accelerators and reducing wasted spend, while also moving Dynamic Resource Allocation for GPUs to general availability.
Kubernetes 1.37 introduces a beta “scale‑to‑zero” feature for GPU workloads and graduates Dynamic Resource Allocation for GPUs to general availability, a move that could dramatically cut idle cloud spending for AI‑heavy applications.
What’s new in Kubernetes 1.37
The release adds a beta GPU scale‑to‑zero controller that automatically reduces pods using GPUs to zero replicas when they are idle. When a pod scales back up, the scheduler instantly provisions the required accelerators, eliminating the need for manual intervention.
Dynamic Resource Allocation (DRA) for GPUs, previously in preview, is now GA. DRA enables pods to request GPU resources at the container level, allowing finer‑grained sharing of a single physical GPU across multiple workloads.
Benefits for cloud cost management
By shrinking idle GPU pods to zero, organizations can avoid paying for unused accelerator time, which has been a major expense for machine‑learning pipelines that run intermittently.
General‑availability DRA also improves utilization, as multiple low‑intensity jobs can now share a single GPU, reducing the total number of devices required in a cluster.
- Automatic pod scaling to zero when no GPU work is pending
- Instant accelerator re‑allocation on pod restart
- Fine‑grained GPU sharing via DRA
- Reduced need for over‑provisioned GPU fleets
How to enable the features
Administrators must enable the GPUScaleToZero feature gate and configure the GPUScaleToZeroController in the kube‑scheduler. For DRA, the DynamicResourceAllocation gate is already on by default, and the ResourceClass objects for GPUs must be defined in the cluster.
Both features rely on the latest device plugin API, so clusters should update their GPU drivers and runtime to the versions bundled with Kubernetes 1.37.
“Scale‑to‑zero for GPUs is a game‑changer for cost‑conscious AI workloads,” said a Kubernetes community maintainer in the release notes.
For detailed installation steps and best‑practice recommendations, see the official Kubernetes 1.37 documentation.
Comments
No comments yet.