Editor’s note: AI21 Labs is a leading global AI lab with a long track record of building foundation models, most notably the Jamba family, and today focuses on specialized LLMs and agent optimization technology. By adopting Google Cloud AI Hypercomputer, AI21 cut high-priority job wait times from 72 hours to 12 and manual scheduling interventions from 20 per week to zero.
At AI21, we build foundation models and agent optimization products that help enterprises run agents at frontier quality, efficiently. Our language models, including the Jamba family, and our agent optimization product suite run demanding production workloads, including our own. We chose Google Cloud AI Hypercomputer to support them at scale.
To keep our model training runs highly utilized, we needed a performant, scalable environment codesigned across infrastructure, orchestration, and consumption models. Our model training runs on one of our shared Google Kubernetes Engine (GKE) clusters, pooling thousands of Google Cloud A3 (powered by NVIDIA H100 Tensor Core GPUs) and A3 Ultra (powered by NVIDIA H200 Tensor Core GPUs) instances, so any team can draw on the full capacity of the fleet rather than being boxed into its own slice. The cluster also trains models and agent-optimization workloads beyond the Jamba family. That approach keeps utilization high, and it makes scheduling hard.
Prior to leveraging GKE for orchestration, we used to negotiate capacity by hand in Slack. If you needed capacity for a training run, you posted in #gpu-resources and hoped for the best.
That worked fine when the cluster had headroom. It stopped working once utilization pinned near 100%, which is where you want a reserved compute fleet to sit.
Over time, every request became a negotiation. Team leads spent their time refereeing compute disputes. Our high-priority jobs — the large, multi-node training runs that need half or more of the cluster at once and serve as the critical path for model projects — could sit blocked for up to 72 hours waiting for enough contiguous capacity to open up.
Scarcity created two distinct problems, and it took us a while to see them as separate. The first was contention: determining who gets compute access next, which we resolved through negotiation. The second was fragmentation: capacity that was technically free but scattered in pieces too small for a large job to use, a bin-packing problem no amount of negotiation could fix.
Sometimes we had plenty of capacity free on paper, but it was scattered across different machines in chunks too small for a larger job to actually land. Without all-or-nothing admission, the cluster could reach a deadlock, with machines holding resources without doing useful work until someone stepped in manually.
It was clear the status quo wasn’t working and we needed something better.
We looked at a few open-source batch schedulers, including Apache YuniKorn, Volcano, and Kueue. YuniKorn didn’t cover all our use cases. And while Volcano had more features, integrating it with our environment would have required replacing core Kubernetes scheduler components.
Kueue won on simplicity and integration. It worked with standard Kubernetes, didn’t require replacing core components, and didn’t force us to rewrite our job specs.
Pairing Kueue with AI Hypercomputer’s flexible, open operations through GKE also contributed to our success. We rely on GKE because it gives us the right level of control for compute-intensive AI work — close access to GPU hardware and drivers, without the overhead of managing raw instances ourselves. Kueue’s native integration with GKE, including with Google Cloud capacity types like Spot VMs and Dynamic Workload Scheduler, meant that once our reserved capacity filled up, the same scheduling logic could reach out to elastic capacity automatically instead of leaving jobs stuck.
Adopting Kueue turned into something bigger than a simple process change. Because we partner with Google Cloud, we have a direct line to a Technical Account Manager, who saw an opportunity to make AI21 a design partner for the Kueue team. This way, we wouldn’t just be a user, but a source of real production requirements that could help shape where the tool went next.
One of the first things to come out of that partnership was a requirements document we shared with the Kueue team, describing behavior we needed that didn’t exist yet: fair admission ordering across teams for multi-node jobs, without the preemption that usually comes bundled with fairness. The Kueue team built it. Admission Fair Sharing (AFS) reorders the admission queue to favor teams that have historically used less capacity without disrupting jobs already running. Around the same time, we also turned on Topology Aware Scheduling. This existing feature made Kueue aware of our physical cluster layout, so it can refuse to admit jobs that won’t fit on a single node rather than placing them with nowhere to run.
For us, that’s what the best-case open-source feedback loop looks like: Real requirements surfaced through production use, fed directly back into Kueue. The Slack #gpu-resources channel is archived now. Every workload, whether a debug pod, a multi-node training run, or an inference deployment, gets queued, prioritized, and scheduled automatically. The results were immediate. Manual interventions dropped from about 20 a week to zero. High-priority jobs that used to wait up to 72 hours now wait 12. Fragmentation across the cluster fell from 15% to 8%, and the “zombie job” problem — workloads partially admitted with nowhere to run — is gone.
None of this changed our total cost. We run our reserved fleet at close to 100% utilization on purpose, so raw spend was never the variable we were optimizing. What changed is where our people’s time goes. Team leads aren’t refereeing compute disputes anymore, and researchers aren’t waiting on replies in Slack. That frees everyone to run more experiments and iterate faster.