cd /news/ai-infrastructure/kubernetes-1-37-garhwal-hpa-scale-to… · home › topics › ai-infrastructure › article
[ARTICLE · art-143159] src=dev.to ↗ pub= topic=ai-infrastructure verified=true sentiment=· neutral

Kubernetes 1.37 Garhwal: HPA Scale-to-Zero, Gang Scheduling, and 18 Removed Kubelet Flags

Kubernetes 1.37 "Garhwal" has been released with 67 enhancements, 16 of which reached Stable, headlined by a HorizontalPodAutoscaler that can scale workloads down to zero replicas and Beta Gang Scheduling (KEP-4671) that starts PodGroups together for distributed AI/ML training jobs. The release also graduates the metrics.k8s.io API to v1 after nine years in beta, stabilizes four Dynamic Resource Allocation features, and removes 18 Kubelet flags as part of the embedded cAdvisor migration, so Kubelets passing removed flags such as --containerd will fail to start. Operators must also move to containerd 2.x, as containerd 1.x is no longer supported.

by read3 min views2 publishedOct 1, 2026

Since August 26, 2026, Kubernetes 1.37 "Garhwal" is available – with 67 enhancements, 16 of which have reached Stable status. The release is named after the Garhwal region in the Indian Himalayas and focuses primarily on consolidation: an API that has been in beta for nine years finally becomes stable; the HorizontalPodAutoscaler can scale workloads down to zero replicas; and Gang Scheduling for AI/ML jobs reaches Beta status. Anyone operating a cluster should prepare for some breaking changes.

Perhaps the most important innovation for operations is the scaled HorizontalPodAutoscaler (HPA). As of Kubernetes 1.37, an HPA can scale workloads down to zero replicas – Beta status, enabled by default. The prerequisite is the use of Object or External metrics, since CPU and memory metrics depend on running Pods. The trick: a queue length exists independently of the workers processing it. The HPA can still read this value even with zero Pods and start new replicas when needed.

The use case is clear: queue consumers, batch jobs, and GPU workloads that are idle between bursts simply consume no resources anymore. A ScaledToZero status on the HPA object reliably distinguishes between "scaled down by the controller" and "manually set to zero".

Distributed training jobs often fail because some Pods are placed while the rest are stuck in a queue – the job does not progress but blocks resources. Gang Scheduling (KEP-4671) solves this problem with the PodGroup concept: the scheduler receives the information that eight Pods must be started together and waits until enough capacity is available for all.

Also newly added is workload-aware Preemption (KEP-5710), which prevents competing workloads from endlessly preempting each other. For platform teams running Ray, JobSet, or LWS, this is the next step toward a production-ready AI platform on Kubernetes.

The metrics.k8s.io API has been in beta since Kubernetes 1.8 – that's nearly nine years. With 1.37, it is stable as v1. The API surface does not change; it is a pure version graduation, but it sends a clear signal: Kubernetes' resource metric infrastructure is production-ready. kubectl top already prefers the v1 version with fallback to v1beta1.

Upgrading to 1.37 requires preparation. The migration of the embedded cAdvisor to a slim cadvisor/lib module (PR #139870) removes 18 Kubelet flags – including --containerd, --containerd-namespace, --boot-id-file, and --enable-load-reader. A Kubelet that passes any of these flags will not start at all with unknown flag. Nodes configured via kubeadm-flags.env or systemd drop-in must be cleaned up before the upgrade.

Also removed: the old IPVS support in kube-proxy finally gives way to nftables, and failCgroupV1 remains enabled by default. Container runtimes must use containerd 2.x – containerd 1.x is no longer supported.

Dynamic Resource Allocation (DRA) receives four Stable graduations in one release – this shows that GPUs, accelerators, and special network adapters are now first-class citizens in Kubernetes' resource model. The standardized DRA attribute resource.kubernetes.io/numaNode allows consistent NUMA topology-dependent placement. Also new in Alpha status is the ulimits configuration per container via the Container.SecurityContext field – a long-awaited feature for databases and highly concurrent workloads.

Pod certificates (PodCertificateRequest) are stable with 1.37. Instead of mounting long-lived Secrets, the Kubelet can request short-lived X.509 certificates and rotate them regularly – a native workload identity system. Also stable: SELinuxMount. Instead of recursively relabeling the entire volume before each Pod start, Kubernetes uses the Linux kernel-native -o context flag and assigns the correct SEL

── more in #ai-infrastructure 4 stories · sorted by recency
── more on @kubernetes 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/kubernetes-1-37-garh…] indexed:0 read:3min 2026-10-01 · —