{"slug": "ai21-achieves-an-83-reduction-in-time-to-start-for-ai-workloads-with-ai", "title": "AI21 achieves an 83% reduction in time-to-start for AI workloads with AI Hypercomputer", "summary": "AI21 Labs cut high-priority AI job wait times from 72 hours to 12 hours — an 83% reduction — and reduced manual scheduling interventions from 20 per week to zero by adopting Google Cloud AI Hypercomputer with the Kueue batch scheduler on Google Kubernetes Engine, according to an AI21 account published by Google Cloud. AI21 runs its Jamba-family model training and agent-optimization workloads on shared GKE clusters pooling thousands of Google Cloud A3 (NVIDIA H100) and A3 Ultra (NVIDIA H200) instances, where near-100% utilization had caused contention and fragmentation that blocked large multi-node jobs. AI21 chose Kueue over Apache YuniKorn and Volcano because it integrated with standard Kubernetes without replacing core scheduler components or rewriting job specs.", "body_md": "**Editor’s note**: AI21 Labs is a leading global AI lab with a long track record of building foundation models, most notably the Jamba family, and today focuses on specialized LLMs and agent optimization technology. By adopting Google Cloud AI Hypercomputer, AI21 cut high-priority job wait times from 72 hours to 12 and manual scheduling interventions from 20 per week to zero. \n\nAt [AI21](https://www.ai21.com/), we build foundation models and agent optimization products that help enterprises run agents at frontier quality, efficiently. Our language models, including the [Jamba](https://www.ai21.com/jamba/) family, and our agent optimization product suite run demanding production workloads, including our own. We chose [Google Cloud AI Hypercomputer](https://cloud.google.com/ai-infrastructure) to support them at scale.\n\nTo keep our model training runs highly utilized, we needed a performant, scalable environment codesigned across infrastructure, orchestration, and consumption models. Our model training runs on one of our shared [Google Kubernetes Engine](https://cloud.google.com/kubernetes-engine) (GKE) clusters, pooling thousands of Google Cloud A3 (powered by NVIDIA H100 Tensor Core GPUs) and A3 Ultra (powered by NVIDIA H200 Tensor Core GPUs) instances, so any team can draw on the full capacity of the fleet rather than being boxed into its own slice. The cluster also trains models and agent-optimization workloads beyond the Jamba family. That approach keeps utilization high, and it makes scheduling hard. \n\nPrior to leveraging GKE for orchestration, we used to negotiate capacity by hand in Slack. If you needed capacity for a training run, you posted in #gpu-resources and hoped for the best.\n\nThat worked fine when the cluster had headroom. It stopped working once utilization pinned near 100%, which is where you want a reserved compute fleet to sit.\n\nOver time, every request became a negotiation. Team leads spent their time refereeing compute disputes. Our high-priority jobs — the large, multi-node training runs that need half or more of the cluster at once and serve as the critical path for model projects — could sit blocked for up to 72 hours waiting for enough contiguous capacity to open up.\n\nScarcity created two distinct problems, and it took us a while to see them as separate. The first was contention: determining who gets compute access next, which we resolved through negotiation. The second was fragmentation: capacity that was technically free but scattered in pieces too small for a large job to use, a bin-packing problem no amount of negotiation could fix.\n\nSometimes we had plenty of capacity free on paper, but it was scattered across different machines in chunks too small for a larger job to actually land. Without all-or-nothing admission, the cluster could reach a deadlock, with machines holding resources without doing useful work until someone stepped in manually.\n\nIt was clear the status quo wasn’t working and we needed something better.\n\nWe looked at a few open-source batch schedulers, including Apache YuniKorn, Volcano, and [Kueue](https://kueue.sigs.k8s.io/). YuniKorn didn’t cover all our use cases. And while Volcano had more features, integrating it with our environment would have required replacing core Kubernetes scheduler components. \n\nKueue won on simplicity and integration. It worked with standard Kubernetes, didn’t require replacing core components, and didn’t force us to rewrite our job specs.\n\nPairing Kueue with AI Hypercomputer’s flexible, open operations through GKE also contributed to our success. We rely on GKE because it gives us the right level of control for compute-intensive AI work — close access to GPU hardware and drivers, without the overhead of managing raw instances ourselves. Kueue’s native integration with GKE, including with Google Cloud capacity types like [Spot VMs](https://cloud.google.com/spot-vms) and [Dynamic Workload Scheduler](https://cloud.google.com/blog/products/compute/introducing-dynamic-workload-scheduler), meant that once our reserved capacity filled up, the same scheduling logic could reach out to elastic capacity automatically instead of leaving jobs stuck.\n\nAdopting Kueue turned into something bigger than a simple process change. Because we partner with Google Cloud, we have a direct line to a Technical Account Manager, who saw an opportunity to make AI21 a design partner for the Kueue team. This way, we wouldn’t just be a user, but a source of real production requirements that could help shape where the tool went next.\n\nOne of the first things to come out of that partnership was a requirements document we shared with the Kueue team, describing behavior we needed that didn’t exist yet: fair admission ordering across teams for multi-node jobs, without the preemption that usually comes bundled with fairness. The Kueue team built it. [Admission Fair Sharing](https://kueue.sigs.k8s.io/docs/concepts/admission_fair_sharing/) (AFS) reorders the admission queue to favor teams that have historically used less capacity without disrupting jobs already running. Around the same time, we also turned on [Topology Aware Scheduling](https://kueue.sigs.k8s.io/docs/concepts/topology_aware_scheduling/). This existing feature made Kueue aware of our physical cluster layout, so it can refuse to admit jobs that won’t fit on a single node rather than placing them with nowhere to run. \n\nFor us, that’s what the best-case open-source feedback loop looks like: Real requirements surfaced through production use, fed directly back into Kueue.\n\nThe Slack #gpu-resources channel is archived now. Every workload, whether a debug pod, a multi-node training run, or an inference deployment, gets queued, prioritized, and scheduled automatically. The results were immediate. Manual interventions dropped from about 20 a week to zero. High-priority jobs that used to wait up to 72 hours now wait 12. Fragmentation across the cluster fell from 15% to 8%, and the “zombie job” problem — workloads partially admitted with nowhere to run — is gone.\n\nNone of this changed our total cost. We run our reserved fleet at close to 100% utilization on purpose, so raw spend was never the variable we were optimizing. What changed is where our people’s time goes. Team leads aren’t refereeing compute disputes anymore, and researchers aren’t waiting on replies in Slack. That frees everyone to run more experiments and iterate faster.", "url": "https://wpnews.pro/news/ai21-achieves-an-83-reduction-in-time-to-start-for-ai-workloads-with-ai", "canonical_source": "https://cloud.google.com/blog/products/containers-kubernetes/ai21-trains-its-models-on-ai-hypercomputer/", "published_at": "2026-10-02 19:00:00+00:00", "updated_at": "2026-10-02 19:39:07.750928+00:00", "lang": "en", "topics": ["ai-infrastructure", "mlops", "large-language-models", "ai-agents"], "entities": ["AI21 Labs", "Google Cloud", "Google Cloud AI Hypercomputer", "Google Kubernetes Engine", "Kueue", "NVIDIA H100", "NVIDIA H200", "Apache YuniKorn"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/ai21-achieves-an-83-reduction-in-time-to-start-for-ai-workloads-with-ai", "markdown": "https://wpnews.pro/news/ai21-achieves-an-83-reduction-in-time-to-start-for-ai-workloads-with-ai.md", "text": "https://wpnews.pro/news/ai21-achieves-an-83-reduction-in-time-to-start-for-ai-workloads-with-ai.txt", "jsonld": "https://wpnews.pro/news/ai21-achieves-an-83-reduction-in-time-to-start-for-ai-workloads-with-ai.jsonld"}}