GPU Job Queueing with Kueue and DRA: Scheduling AI Workloads Without Idle Capacity Cast AI's 2026 State of Kubernetes Optimization Report found average GPU utilization across its fleet of 23,000 production clusters is just 5%, with AKS at 2%, EKS at 5%, and GKE at 6%, while the best observed fleet — a 136-node H200 LLM inference cluster — sustains 49%. The report attributes the gap to scheduling and sharing rather than hardware, and points to Kueue, a Kubernetes-native admission controller that holds GPU jobs until capacity and quota are confirmed available, plus Dynamic Resource Allocation (GA in Kubernetes 1.34) as fixes; Kueue 1.1+ adds JobSet and LeaderWorkerSet gang scheduling, and the current stable release is v1.3.0. Cast AI says teams running MIG partitioning with DRA on H100 clusters cut required GPU count from 100 GPUs to 25–75 GPUs while serving equivalent workloads. Kueue is a Kubernetes-native job queueing controller: instead of admitting a GPU job that cannot be scheduled and leaving it pending indefinitely, it holds the job until the quota and capacity it needs are actually available, then admits it. Dynamic Resource Allocation is the Kubernetes API that lets a workload describe the device it needs – a GPU with particular memory or a particular sharing mode – rather than asking for a whole card. Together they address the two reasons GPU capacity sits idle at an average of 5%: jobs that cannot be placed, and jobs that hold more device than they use. Key takeaways - Kueue is a Kubernetes-native admission controller that holds GPU jobs in a queue until capacity and quota are confirmed available, eliminating zombie Pending pods. - DRA GA in Kubernetes 1.34 https://kubernetes.io/blog/2025/09/01/kubernetes-v1-34-dra-updates/ lets workloads describe exact hardware requirements – minimum VRAM, compute capability – instead of requesting opaque integer GPU counts. - Kueue 1.1+ adds first-class JobSet and LeaderWorkerSet support for gang scheduling without a second scheduler binary; the current stable release is v1.3.0. - According to Cast AI’s 2026 State of Kubernetes Optimization Report, teams running MIG partitioning with DRA on H100 clusters reduced required GPU count from 100 GPUs to 25–75 GPUs while serving equivalent workloads. - Cast AI’s autoscaler reads ResourceClaim attributes and provisions the exact matching GPU instance type automatically, completing the loop from queue to node. Cast AI’s automation platform measures utilization across 23,000 production clusters. The fleet data from early 2026 is clear: average GPU utilization is 5%. AKS clusters average 2%. EKS averages 5%. GKE reaches 6%. The best observed fleet, a 136-node H200 LLM inference cluster, sustains 49%. That 10x gap has almost nothing to do with hardware. It is a scheduling and sharing problem. Kueue GPU scheduling turns the Kubernetes scheduler into a quota-aware admission controller; Dynamic Resource Allocation DRA tells it exactly which device each job needs. Together they address both structural causes of 5% utilization: jobs that cannot be placed, and jobs that hold more device than they use. What problem queueing solves for GPU workloads The default Kubernetes scheduler is a placement engine. It does not manage quota, enforce fairness, or hold jobs when capacity is unavailable. When a GPU pod targets a node that lacks capacity, kube-scheduler marks it Pending and moves on. That pod then sits consuming no GPU but holding a scheduling reservation, in a state that operators cannot inspect, reorder, or preempt. For broader context on GPU waste patterns in Kubernetes, see Kubernetes GPU Optimization: How to Cut GPU Waste Without Slowing Workloads https://cast.ai/blog/kubernetes-gpu-optimization/ . Before going further: not all GPU workloads queue the same way. Batch training jobs PyTorchJob, fine-tuning runs need guaranteed gang scheduling, every worker must start together or none should start. Inference servers such as vLLM and TGI rely on dynamic request batching rather than job-level queuing. This post focuses on batch training and job queuing. For inference autoscaling and GPU sharing for serving workloads, see GPU Sharing in Kubernetes: How to Cut Costs with Cast AI https://cast.ai/blog/gpu-sharing-kubernetes-cost-optimization/ . The zombie Pending pod pattern Three failure modes result from this design. First, Pending pods create an invisible queue: platform engineers cannot determine whether a job waits because all GPUs are busy, a previous job is stuck, or no node has the required GPU type. Second, for distributed training, partial scheduling destroys efficiency. Some workers start while others wait, burning GPU time on processes that cannot make forward progress without all ranks present. Third, without quota enforcement, a single team can monopolize every cluster GPU for hours or days, blocking higher-priority workloads entirely. Kueue solves this by sitting above kube-scheduler as an admission controller. Jobs are created in a Suspended state. Kueue holds each Workload in an observable queue, evaluates quota and capacity, and flips spec.suspend=false to release it to kube-scheduler for pod placement. No job ever becomes a zombie Pending pod. The queue is visible, orderable, and preemptable. How Kueue works: ClusterQueue, LocalQueue, ResourceFlavor, admission Kueue’s architecture centers on four objects. Understanding each one clarifies where the configuration levers sit for multi-tenant GPU clusters. Install Kueue first: kubectl apply --server-side -f https://github.com/kubernetes-sigs/kueue/releases/download/v1.3.0/manifests.yaml Alternatively, use the Helm chart: helm install kueue oci://registry.k8s.io/kueue-charts/kueue --version v1.3.0 -n kueue-system --create-namespace . Once CRDs are installed, configure queues and quotas as follows. The four core objects ResourceFlavor describes a pool of like hardware. It maps to a node group via node labels and tells Kueue which GPU model and billing model a set of physical nodes represents. Separate flavors for a100-ondemand and h100-spot let burst jobs consume spot-GPU quota before touching on-demand quota. ClusterQueue is the cluster-scoped resource pool governor. It defines nominalQuota the team’s guaranteed allocation , borrowingLimit how much extra quota it pulls from idle cohort-mates , and lendingLimit how much of its own unused quota it shares . Kueue v1.3.0 current stable uses the v1beta2 API for all ClusterQueue resources, and multiple resource groups let you track CPU, memory, and nvidia.com/gpu together per flavor. LocalQueue is the namespaced handle teams interact with directly. It references a ClusterQueue and acts as the user-facing submission gateway, so application teams submit jobs without needing cluster-wide permissions. Workload is Kueue’s internal admission unit, wrapping any supported job type: batch/v1 Job , JobSet, PyTorchJob, TFJob, RayJob, MPIJob, and plain Pods. Kueue evaluates the Workload’s total resource sum against ClusterQueue quota, assigns a ResourceFlavor, and either admits or holds it atomically. Start with a ResourceFlavor that maps your H100 node pool to a named flavor — this is the first object in the Kueue hierarchy and is required before ClusterQueue will admit any work: apiVersion: kueue.x-k8s.io/v1beta1 kind: ResourceFlavor metadata: name: nvidia-h100 spec: nodeLabels: cloud.google.com/gke-accelerator: nvidia-tesla-h100 or node.kubernetes.io/instance-type: p5.48xlarge for AWS tolerations: - key: nvidia.com/gpu operator: Exists effect: NoSchedule ResourceFlavor maps a hardware type here: H100 nodes to a named flavor. Your ClusterQueue’s resourceGroups should reference this flavor name. A minimal ClusterQueue + LocalQueue pairing to enforce GPU quota per team: apiVersion: kueue.x-k8s.io/v1beta2 kind: ClusterQueue metadata: name: team-alpha-cq spec: cohort: gpu-pool resourceGroups: - coveredResources: "nvidia.com/gpu" flavors: - name: nvidia-h100 resources: - name: "nvidia.com/gpu" nominalQuota: 8 borrowingLimit: 4 --- apiVersion: kueue.x-k8s.io/v1beta2 kind: LocalQueue metadata: name: training-queue namespace: team-alpha spec: clusterQueue: team-alpha-cq Admission and queueing strategies Kueue supports two queueing strategies. StrictFIFO processes workloads in priority-then-creation-time order; a large job at the front blocks smaller jobs even when capacity exists for them. BestEffortFIFO allows newer, lower-resource jobs to slip past a blocked older job when capacity permits, reducing average wait time significantly for most AI workloads. Kueue also emits Prometheus metrics for queue depth, admission latency, and eviction latency, making GPU capacity planning observable for the first time on many clusters. What Dynamic Resource Allocation changes The legacy device plugin model is count-based. A node advertises nvidia.com/gpu: 8 , and pods request integers. There is no attribute visibility, no cross-pod sharing negotiation, and no topology awareness. This forces a binary choice: claim a whole GPU or get nothing. Consequently, a fine-tuning job needing 12GB of VRAM claims a full 80GB H100, leaving 68GB idle. DRA replaces opaque integer counts with structured attribute-based claims. It graduated to GA in Kubernetes 1.34 https://kubernetes.io/blog/2025/09/01/kubernetes-v1-34-dra-updates/ see also KEP-3063 https://github.com/kubernetes/enhancements/tree/master/keps/sig-node/3063-dynamic-resource-allocation and is enabled by default from 1.34 onward. DRA introduces ResourceSlice driver-published inventory with structured attributes including GPU model, VRAM, architecture, and CUDA compute capability , DeviceClass admin-defined device category , and ResourceClaim a workload’s specific hardware request . For a practical deployment walkthrough, see Deploying GPU Workloads with Dynamic Resource Allocation https://cast.ai/blog/deploying-gpu-workload-with-dynamic-resource-allocation/ . MIG slices, MPS, and the scheduling flow MIG integration is where DRA shows its practical value. The NVIDIA DRA driver creates a mig.nvidia.com DeviceClass. Admins define custom DeviceClasses like mig-3g.40gb , and workloads reference specific MIG profiles in their ResourceClaims rather than relying on static node labels. For teams already using MIG and time-slicing, GPU Sharing in Kubernetes: How to Cut Costs with Cast AI https://cast.ai/blog/gpu-sharing-kubernetes-cost-optimization/ covers complementary partitioning techniques in detail. One boundary to understand: DRA selects the MIG slice based on claim attributes, while MIG profiles define the available slice sizes. On H100, available slice sizes are roughly 10GB, 20GB, 40GB, and 80GB. A 12GB fine-tuning job claims a 20GB MIG slice, leaving approximately 8GB stranded within that slice. If workloads cannot tolerate MIG’s hard per-partition isolation, NVIDIA’s Multi-Process Service MPS is an alternative: MPS shares a single GPU context across multiple processes without hardware partitioning, reducing fragmentation at the cost of less strict isolation between tenants. When to use MIG vs MPS vs full-GPU allocation NVLink is disabled between MIG instances – peer-to-peer GPU memory access is unavailable within a MIG-partitioned card. MIG works well for independent inference serving, eval workloads, and single-GPU fine-tuning where each job fits in one partition. For NCCL-based distributed training PyTorch DDP, Megatron-LM , do not use MIG: allocate full GPUs so NVLink remains available for high-bandwidth collective operations. For shared inference workloads where you want time-multiplexed access without the NVLink penalty, MPS Multi-Process Service provides time-shared access to a single GPU context without disabling NVLink. A ResourceClaim referencing a MIG slice by attribute: apiVersion: resource.k8s.io/v1 kind: ResourceClaim metadata: name: h100-mig-slice namespace: team-alpha spec: devices: requests: - name: gpu deviceClassName: mig-3g.40gb selectors: - cel: Attribute key — verify against your NVIDIA GPU Operator DRA driver version before applying expression: device.attributes "memory" .isGreaterThan quantity "39Gi" The DRA scheduling flow runs in five steps: the driver publishes ResourceSlices, the pod creates a ResourceClaim, kube-scheduler’s DynamicResources plugin matches claim requirements against ResourceSlices cluster-wide, the allocation result writes to ResourceClaim.status , and the container runtime injects the device. Error visibility also improves: the ResourceClaim object shows “no device matches selector” when a claim cannot be satisfied, so engineers do not need to dig through Pending pod logs. Kueue and DRA together DRA and Kueue solve complementary problems but need explicit integration to cooperate on quota accounting. Kueue 1.1+ introduced the ResourceClaimTemplate path beta . Kueue 1.2+ added the extended resource path beta , graduating KueueDRAIntegrationExtendedResource to enabled-by-default. Two integration paths and one timing caveat With the ResourceClaimTemplate path , pods explicitly reference a ResourceClaimTemplate and Kueue maps DeviceClass names to quota resource names via deviceClassMappings in Kueue Configuration. With the extended resource path , pods use the traditional resources.requests syntax for example, nvidia.com/gpu: 1 and the DeviceClass carries an extendedResourceName field. Kube-scheduler auto-creates the ResourceClaim, and Kueue detects the DeviceClass to avoid double-counting quota. Start with the extended resource path if your pods already use resources.requests: nvidia.com/gpu: 1 – Kueue detects DRA usage automatically and avoids double-counting. Use the ResourceClaimTemplate path when you need per-pod control over specific MIG profiles, NVLink clique membership, or VRAM thresholds. One timing caveat matters in production: Kueue admits the workload the quota check before kube-scheduler allocates the actual device. If cluster state changes between those two steps, the scheduler may fail to allocate. WaitForPodsReady handles this by evicting workloads that fail to become ready within a configured timeout, returning their quota to the queue for retry. The combined effect is precise: Kueue holds a job until both quota and capacity are confirmed. DRA then selects the right device—including which MIG slice, which NVLink clique, and which GPU has sufficient VRAM. Jobs no longer pile up in Pending, and they no longer claim more device than they need. Quota, fairness and preemption across teams Multi-team GPU clusters need more than a single shared queue. Kueue’s cohort model gives teams guaranteed capacity while letting idle GPUs flow to whoever has pending work. Cohort borrowing and reclaim Multiple ClusterQueues grouped into a cohort can borrow each other’s unused nominalQuota . A team with 8 guaranteed GPUs can burst to 16 if other teams’ GPUs sit idle, subject to their borrowingLimit . When the owning team submits new work, reclaimWithinCohort: Any preempts the borrower’s jobs and returns the quota immediately. Preemption operates at three scopes: withinClusterQueue evicts lower-priority jobs in the same queue, reclaimWithinCohort evicts cohort-mate jobs running above their nominalQuota , and borrowWithinCohort allows preemption of lower-priority cohort-mate workloads to borrow quota. These policies together eliminate the starvation scenarios that plague unmanaged multi-team GPU clusters. Gang scheduling for distributed training Distributed training frameworks like PyTorch DDP use NCCL collective operations that require all ranks to be present simultaneously. If one worker pod is missing, the collective blocks. Partial scheduling of a distributed job therefore wastes every GPU that is already running. Kueue 1.1+ and JobSet Gang admission happens at the quota level: either the entire Workload, including all PodSets, is admitted or none of it is. Kueue 1.1+ provides first-class support for JobSet and LeaderWorkerSet, giving pod-level gang coordination on top of the default scheduler. JobSet manages coordinated groups of Jobs with shared failure, restart, and success semantics, eliminating the need for a second scheduler binary in most AI batch use cases. For PyTorchJob workloads, add the label kueue.x-k8s.io/queue-name: