{"slug": "gpu-job-queueing-with-kueue-and-dra-scheduling-ai-workloads-without-idle", "title": "GPU Job Queueing with Kueue and DRA: Scheduling AI Workloads Without Idle Capacity", "summary": "Cast AI's 2026 State of Kubernetes Optimization Report found average GPU utilization across its fleet of 23,000 production clusters is just 5%, with AKS at 2%, EKS at 5%, and GKE at 6%, while the best observed fleet — a 136-node H200 LLM inference cluster — sustains 49%. The report attributes the gap to scheduling and sharing rather than hardware, and points to Kueue, a Kubernetes-native admission controller that holds GPU jobs until capacity and quota are confirmed available, plus Dynamic Resource Allocation (GA in Kubernetes 1.34) as fixes; Kueue 1.1+ adds JobSet and LeaderWorkerSet gang scheduling, and the current stable release is v1.3.0. Cast AI says teams running MIG partitioning with DRA on H100 clusters cut required GPU count from 100 GPUs to 25–75 GPUs while serving equivalent workloads.", "body_md": "Kueue is a Kubernetes-native job queueing controller: instead of admitting a GPU job that cannot be scheduled and leaving it pending indefinitely, it holds the job until the quota and capacity it needs are actually available, then admits it. Dynamic Resource Allocation is the Kubernetes API that lets a workload describe the device it needs – a GPU with particular memory or a particular sharing mode – rather than asking for a whole card. Together they address the two reasons GPU capacity sits idle at an average of 5%: jobs that cannot be placed, and jobs that hold more device than they use.\n\n## Key takeaways\n\n- Kueue is a Kubernetes-native admission controller that holds GPU jobs in a queue until capacity and quota are confirmed available, eliminating zombie Pending pods.\n- DRA ([GA in Kubernetes 1.34](https://kubernetes.io/blog/2025/09/01/kubernetes-v1-34-dra-updates/) ) lets workloads describe exact hardware requirements – minimum VRAM, compute capability – instead of requesting opaque integer GPU counts.\n- Kueue 1.1+ adds first-class JobSet and LeaderWorkerSet support for gang scheduling without a second scheduler binary; the current stable release is v1.3.0.\n- According to Cast AI’s 2026 State of Kubernetes Optimization Report, teams running MIG partitioning with DRA on H100 clusters reduced required GPU count from 100 GPUs to 25–75 GPUs while serving equivalent workloads.\n- Cast AI’s autoscaler reads ResourceClaim attributes and provisions the exact matching GPU instance type automatically, completing the loop from queue to node.\n\nCast AI’s automation platform measures utilization across 23,000 production clusters. The fleet data from early 2026 is clear: average GPU utilization is 5%. AKS clusters average 2%. EKS averages 5%. GKE reaches 6%. The best observed fleet, a 136-node H200 LLM inference cluster, sustains 49%. That 10x gap has almost nothing to do with hardware. It is a scheduling and sharing problem. Kueue GPU scheduling turns the Kubernetes scheduler into a quota-aware admission controller; Dynamic Resource Allocation (DRA) tells it exactly which device each job needs. Together they address both structural causes of 5% utilization: jobs that cannot be placed, and jobs that hold more device than they use.\n\n## What problem queueing solves for GPU workloads\n\nThe default Kubernetes scheduler is a placement engine. It does not manage quota, enforce fairness, or hold jobs when capacity is unavailable. When a GPU pod targets a node that lacks capacity, kube-scheduler marks it Pending and moves on. That pod then sits consuming no GPU but holding a scheduling reservation, in a state that operators cannot inspect, reorder, or preempt. For broader context on GPU waste patterns in Kubernetes, see [Kubernetes GPU Optimization: How to Cut GPU Waste Without Slowing Workloads](https://cast.ai/blog/kubernetes-gpu-optimization/).\n\nBefore going further: not all GPU workloads queue the same way. Batch training jobs (PyTorchJob, fine-tuning runs) need guaranteed gang scheduling, every worker must start together or none should start. Inference servers such as vLLM and TGI rely on dynamic request batching rather than job-level queuing. This post focuses on batch training and job queuing. For inference autoscaling and GPU sharing for serving workloads, see [GPU Sharing in Kubernetes: How to Cut Costs with Cast AI](https://cast.ai/blog/gpu-sharing-kubernetes-cost-optimization/).\n\n### The zombie Pending pod pattern\n\nThree failure modes result from this design. First, Pending pods create an invisible queue: platform engineers cannot determine whether a job waits because all GPUs are busy, a previous job is stuck, or no node has the required GPU type. Second, for distributed training, partial scheduling destroys efficiency. Some workers start while others wait, burning GPU time on processes that cannot make forward progress without all ranks present. Third, without quota enforcement, a single team can monopolize every cluster GPU for hours or days, blocking higher-priority workloads entirely.\n\nKueue solves this by sitting above kube-scheduler as an admission controller. Jobs are created in a Suspended state. Kueue holds each Workload in an observable queue, evaluates quota and capacity, and flips `spec.suspend=false` to release it to kube-scheduler for pod placement. No job ever becomes a zombie Pending pod. The queue is visible, orderable, and preemptable.\n\n## How Kueue works: ClusterQueue, LocalQueue, ResourceFlavor, admission\n\nKueue’s architecture centers on four objects. Understanding each one clarifies where the configuration levers sit for multi-tenant GPU clusters.\n\nInstall Kueue first:\n\n```\nkubectl apply --server-side -f https://github.com/kubernetes-sigs/kueue/releases/download/v1.3.0/manifests.yaml\n```\n\nAlternatively, use the Helm chart: `helm install kueue oci://registry.k8s.io/kueue-charts/kueue --version v1.3.0 -n kueue-system --create-namespace`. Once CRDs are installed, configure queues and quotas as follows.\n\n### The four core objects\n\n**ResourceFlavor** describes a pool of like hardware. It maps to a node group via node labels and tells Kueue which GPU model and billing model a set of physical nodes represents. Separate flavors for `a100-ondemand` and `h100-spot` let burst jobs consume spot-GPU quota before touching on-demand quota.\n\n**ClusterQueue** is the cluster-scoped resource pool governor. It defines `nominalQuota` (the team’s guaranteed allocation), `borrowingLimit` (how much extra quota it pulls from idle cohort-mates), and `lendingLimit` (how much of its own unused quota it shares). Kueue v1.3.0 (current stable) uses the `v1beta2` API for all ClusterQueue resources, and multiple resource groups let you track CPU, memory, and `nvidia.com/gpu` together per flavor.\n\n**LocalQueue** is the namespaced handle teams interact with directly. It references a ClusterQueue and acts as the user-facing submission gateway, so application teams submit jobs without needing cluster-wide permissions. **Workload** is Kueue’s internal admission unit, wrapping any supported job type: `batch/v1 Job`, JobSet, PyTorchJob, TFJob, RayJob, MPIJob, and plain Pods. Kueue evaluates the Workload’s total resource sum against ClusterQueue quota, assigns a ResourceFlavor, and either admits or holds it atomically.\n\nStart with a ResourceFlavor that maps your H100 node pool to a named flavor — this is the first object in the Kueue hierarchy and is required before ClusterQueue will admit any work:\n\n```\napiVersion: kueue.x-k8s.io/v1beta1\nkind: ResourceFlavor\nmetadata:\n  name: nvidia-h100\nspec:\n  nodeLabels:\n    cloud.google.com/gke-accelerator: nvidia-tesla-h100  # or node.kubernetes.io/instance-type: p5.48xlarge for AWS\n  tolerations:\n  - key: nvidia.com/gpu\n    operator: Exists\n    effect: NoSchedule\n```\n\nResourceFlavor maps a hardware type (here: H100 nodes) to a named flavor. Your ClusterQueue’s `resourceGroups` should reference this flavor name.\n\nA minimal ClusterQueue + LocalQueue pairing to enforce GPU quota per team:\n\n```\napiVersion: kueue.x-k8s.io/v1beta2\nkind: ClusterQueue\nmetadata:\n  name: team-alpha-cq\nspec:\n  cohort: gpu-pool\n  resourceGroups:\n  - coveredResources: [\"nvidia.com/gpu\"]\n    flavors:\n    - name: nvidia-h100\n      resources:\n      - name: \"nvidia.com/gpu\"\n        nominalQuota: 8\n        borrowingLimit: 4\n---\napiVersion: kueue.x-k8s.io/v1beta2\nkind: LocalQueue\nmetadata:\n  name: training-queue\n  namespace: team-alpha\nspec:\n  clusterQueue: team-alpha-cq\n```\n\n### Admission and queueing strategies\n\nKueue supports two queueing strategies. **StrictFIFO** processes workloads in priority-then-creation-time order; a large job at the front blocks smaller jobs even when capacity exists for them. **BestEffortFIFO** allows newer, lower-resource jobs to slip past a blocked older job when capacity permits, reducing average wait time significantly for most AI workloads. Kueue also emits Prometheus metrics for queue depth, admission latency, and eviction latency, making GPU capacity planning observable for the first time on many clusters.\n\n## What Dynamic Resource Allocation changes\n\nThe legacy device plugin model is count-based. A node advertises `nvidia.com/gpu: 8`, and pods request integers. There is no attribute visibility, no cross-pod sharing negotiation, and no topology awareness. This forces a binary choice: claim a whole GPU or get nothing. Consequently, a fine-tuning job needing 12GB of VRAM claims a full 80GB H100, leaving 68GB idle.\n\nDRA replaces opaque integer counts with structured attribute-based claims. It [graduated to GA in Kubernetes 1.34](https://kubernetes.io/blog/2025/09/01/kubernetes-v1-34-dra-updates/) (see also [KEP-3063](https://github.com/kubernetes/enhancements/tree/master/keps/sig-node/3063-dynamic-resource-allocation)) and is enabled by default from 1.34 onward. DRA introduces **ResourceSlice** (driver-published inventory with structured attributes including GPU model, VRAM, architecture, and CUDA compute capability), **DeviceClass** (admin-defined device category), and **ResourceClaim** (a workload’s specific hardware request). For a practical deployment walkthrough, see [Deploying GPU Workloads with Dynamic Resource Allocation](https://cast.ai/blog/deploying-gpu-workload-with-dynamic-resource-allocation/).\n\n### MIG slices, MPS, and the scheduling flow\n\nMIG integration is where DRA shows its practical value. The NVIDIA DRA driver creates a `mig.nvidia.com` DeviceClass. Admins define custom DeviceClasses like `mig-3g.40gb`, and workloads reference specific MIG profiles in their ResourceClaims rather than relying on static node labels. For teams already using MIG and time-slicing, [GPU Sharing in Kubernetes: How to Cut Costs with Cast AI](https://cast.ai/blog/gpu-sharing-kubernetes-cost-optimization/) covers complementary partitioning techniques in detail.\n\nOne boundary to understand: DRA selects the MIG slice based on claim attributes, while MIG profiles define the available slice sizes. On H100, available slice sizes are roughly 10GB, 20GB, 40GB, and 80GB. A 12GB fine-tuning job claims a 20GB MIG slice, leaving approximately 8GB stranded within that slice. If workloads cannot tolerate MIG’s hard per-partition isolation, NVIDIA’s Multi-Process Service (MPS) is an alternative: MPS shares a single GPU context across multiple processes without hardware partitioning, reducing fragmentation at the cost of less strict isolation between tenants.\n\n#### When to use MIG vs MPS vs full-GPU allocation\n\nNVLink is disabled between MIG instances – peer-to-peer GPU memory access is unavailable within a MIG-partitioned card. MIG works well for independent inference serving, eval workloads, and single-GPU fine-tuning where each job fits in one partition. For NCCL-based distributed training (PyTorch DDP, Megatron-LM), do not use MIG: allocate full GPUs so NVLink remains available for high-bandwidth collective operations. For shared inference workloads where you want time-multiplexed access without the NVLink penalty, MPS (Multi-Process Service) provides time-shared access to a single GPU context without disabling NVLink.\n\nA ResourceClaim referencing a MIG slice by attribute:\n\n```\napiVersion: resource.k8s.io/v1\nkind: ResourceClaim\nmetadata:\n  name: h100-mig-slice\n  namespace: team-alpha\nspec:\n  devices:\n    requests:\n    - name: gpu\n      deviceClassName: mig-3g.40gb\n      selectors:\n      - cel:\n          # Attribute key — verify against your NVIDIA GPU Operator DRA driver version before applying\n          expression: device.attributes[\"memory\"].isGreaterThan(quantity(\"39Gi\"))\n```\n\nThe DRA scheduling flow runs in five steps: the driver publishes ResourceSlices, the pod creates a ResourceClaim, kube-scheduler’s DynamicResources plugin matches claim requirements against ResourceSlices cluster-wide, the allocation result writes to `ResourceClaim.status`, and the container runtime injects the device. Error visibility also improves: the ResourceClaim object shows “no device matches selector” when a claim cannot be satisfied, so engineers do not need to dig through Pending pod logs.\n\n## Kueue and DRA together\n\nDRA and Kueue solve complementary problems but need explicit integration to cooperate on quota accounting. Kueue 1.1+ introduced the ResourceClaimTemplate path (beta). Kueue 1.2+ added the extended resource path (beta), graduating `KueueDRAIntegrationExtendedResource` to enabled-by-default.\n\n### Two integration paths and one timing caveat\n\nWith the **ResourceClaimTemplate path**, pods explicitly reference a ResourceClaimTemplate and Kueue maps DeviceClass names to quota resource names via `deviceClassMappings` in Kueue Configuration. With the **extended resource path**, pods use the traditional `resources.requests` syntax (for example, `nvidia.com/gpu: 1`) and the DeviceClass carries an `extendedResourceName` field. Kube-scheduler auto-creates the ResourceClaim, and Kueue detects the DeviceClass to avoid double-counting quota.\n\nStart with the extended resource path if your pods already use `resources.requests: nvidia.com/gpu: 1` – Kueue detects DRA usage automatically and avoids double-counting. Use the ResourceClaimTemplate path when you need per-pod control over specific MIG profiles, NVLink clique membership, or VRAM thresholds.\n\nOne timing caveat matters in production: Kueue admits the workload (the quota check) before kube-scheduler allocates the actual device. If cluster state changes between those two steps, the scheduler may fail to allocate. **WaitForPodsReady** handles this by evicting workloads that fail to become ready within a configured timeout, returning their quota to the queue for retry.\n\nThe combined effect is precise: Kueue holds a job until both quota and capacity are confirmed. DRA then selects the right device—including which MIG slice, which NVLink clique, and which GPU has sufficient VRAM. Jobs no longer pile up in Pending, and they no longer claim more device than they need.\n\n## Quota, fairness and preemption across teams\n\nMulti-team GPU clusters need more than a single shared queue. Kueue’s cohort model gives teams guaranteed capacity while letting idle GPUs flow to whoever has pending work.\n\n### Cohort borrowing and reclaim\n\nMultiple ClusterQueues grouped into a cohort can borrow each other’s unused `nominalQuota`. A team with 8 guaranteed GPUs can burst to 16 if other teams’ GPUs sit idle, subject to their `borrowingLimit`. When the owning team submits new work, `reclaimWithinCohort: Any` preempts the borrower’s jobs and returns the quota immediately. Preemption operates at three scopes: `withinClusterQueue` evicts lower-priority jobs in the same queue, `reclaimWithinCohort` evicts cohort-mate jobs running above their `nominalQuota`, and `borrowWithinCohort` allows preemption of lower-priority cohort-mate workloads to borrow quota. These policies together eliminate the starvation scenarios that plague unmanaged multi-team GPU clusters.\n\n## Gang scheduling for distributed training\n\nDistributed training frameworks like PyTorch DDP use NCCL collective operations that require all ranks to be present simultaneously. If one worker pod is missing, the collective blocks. Partial scheduling of a distributed job therefore wastes every GPU that is already running.\n\n### Kueue 1.1+ and JobSet\n\nGang admission happens at the quota level: either the entire Workload, including all PodSets, is admitted or none of it is. Kueue 1.1+ provides first-class support for JobSet and LeaderWorkerSet, giving pod-level gang coordination on top of the default scheduler. JobSet manages coordinated groups of Jobs with shared failure, restart, and success semantics, eliminating the need for a second scheduler binary in most AI batch use cases.\n\nFor PyTorchJob workloads, add the label `kueue.x-k8s.io/queue-name: <your-LocalQueue>` to the PyTorchJob manifest. Kueue evaluates total resources across the master and all workers as a single Workload, admitting or holding them together:\n\n```\napiVersion: kubeflow.org/v1\nkind: PyTorchJob\nmetadata:\n  name: llm-finetune\n  namespace: team-alpha\n  labels:\n    kueue.x-k8s.io/queue-name: training-queue\nspec:\n  pytorchReplicaSpecs:\n    Master:\n      replicas: 1\n      template:\n        spec:\n          containers:\n          - name: pytorch\n            image: pytorch/pytorch:2.3.0-cuda12.1-cudnn8-runtime\n            resources:\n              limits:\n                nvidia.com/gpu: \"1\"\n    Worker:\n      replicas: 7\n      template:\n        spec:\n          containers:\n          - name: pytorch\n            image: pytorch/pytorch:2.3.0-cuda12.1-cudnn8-runtime\n            resources:\n              limits:\n                nvidia.com/gpu: \"1\"\n```\n\nRayJob gang scheduling uses the same annotation. Add `kueue.x-k8s.io/queue-name: <your-LocalQueue>` to the RayJob `metadata.labels`; Kueue then creates a ProvisioningRequest specifying GPU requirements, the cluster autoscaler adds nodes, and Kueue admits the full RayJob once capacity exists. The annotation itself is a three-line addition identical to the PyTorchJob pattern above.\n\n## Kueue vs Volcano vs the default scheduler\n\nThree options exist for GPU scheduling on Kubernetes. Each fits a different operational profile.\n\n| Feature | Default Scheduler | Kueue | Volcano | \n|---|---|---|---|\n| Scheduling model | Pod placement only; no queue | Admission controller + queue on top of kube-scheduler | Replacement batch scheduler with its own queue primitives | \n| GPU quota per team | None (ResourceQuota only, no borrowing) | nominalQuota, borrowingLimit, lendingLimit per ClusterQueue | Queue weight-based fair-share | \n| Gang scheduling | None | Atomic admission via Workload; gang coordination via JobSet (Kueue 1.1+) | Native PodGroup gang scheduling at scheduler level | \n| Preemption | PriorityClass-based pod eviction | reclaimWithinCohort, borrowWithinCohort, withinClusterQueue policies | Job-level preemption via queue priority | \n| DRA support | Yes (DynamicResources plugin) | Yes: ResourceClaimTemplate path (Kueue 1.1+) and extended resource path (Kueue 1.2+) | Partial; newer and less documented | \n| Multi-tenant fairness | None | Cohort borrowing, Fair Sharing, and preemption policies | Hierarchical queues with fair-share scheduling | \n| Job framework support | Any pod-creating resource | Job, JobSet, PyTorchJob, TFJob, RayJob, MPI, XGBoost, LeaderWorkerSet | TensorFlow, PyTorch, Spark, MPI, Flink, Argo, Ray, PaddlePaddle | \n| Operational complexity | None (built-in) | Low: CRD-based, kubectl-native, no second scheduler binary | Medium: second scheduler binary, PodGroup CRDs, separate queue config | \n| CNCF status | Core Kubernetes | CNCF Incubating (kubernetes-sigs) | CNCF Graduated | \n| Best fit | Single-pod GPU inference, dev clusters | Multi-tenant batch, LLM fine-tuning queues, eval sweeps, mixed CPU+GPU batch | HPC/MPI workloads, Slurm-trained operators, large-scale distributed training with deep gang semantics | \n| Kubernetes baseline | 1.34+ (DRA GA) | Current stable: Kueue v1.3.0 | v1.11+ | \n\nThe practical decision line: use Kueue with JobSet for multi-tenant GPU batch, LLM fine-tuning queues, and agent evaluation sweeps. Add Volcano only when you specifically need MPI-style gang semantics at the pod-placement level, or when your teams expect Volcano’s PodGroup primitives. Kueue 1.1+ plus JobSet closes most of the gap for AI workloads without a second scheduler binary to maintain.\n\n## What this is worth\n\nA single idle H100 on AWS p5 on-demand costs approximately $8,850 per GPU per month. Teams with Reserved Instances or Savings Plans typically pay 30–40% less – adjust the waste calculation to your actual blended rate. At 5% average utilization, 95% of that spend produces no work. A team running a 20-GPU cluster wastes roughly $168,750 per month: 20 GPUs × $8,850 × 95% idle = $168,750 in wasted capacity.\n\n### The before and after\n\nStart with the numbers. $168,750 per month in wasted GPU capacity. According to Cast AI’s 2026 State of Kubernetes Optimization Report, teams that deployed MIG partitioning with DRA on H100 clusters reduced required GPU count from 100 GPUs down to 25–75 GPUs while serving equivalent workloads. Same throughput, one-quarter to three-quarters of the hardware.\n\nWhat changed between those two states: before Kueue and DRA, 100 physical GPUs run 100 jobs simultaneously, one job per whole-card allocation. A 12GB fine-tuning job occupies an 80GB H100; 68GB sits idle. A distributed training job with nine workers starts eight and blocks on the ninth. After: Kueue enforces quota per team, so no team over-claims. DRA selects the right MIG slice per job, so each card runs multiple workloads at once. Idle slices flow immediately to queued work.\n\n### How Cast AI closes the autoscaling loop\n\nKueue and DRA handle scheduling and device allocation. However, when a ResourceClaim cannot be satisfied because no matching node exists, those tools need an autoscaler that understands device attributes. Cast AI’s autoscaler reads ResourceClaim attributes directly, simulates the same DRA allocation checks the Kubernetes scheduler performs, identifies the cheapest instance type that satisfies the workload’s requirements (architecture, VRAM, CUDA capability), and provisions it. When the node joins and the NVIDIA DRA driver publishes a ResourceSlice, the scheduler completes the allocation without custom node labels or workarounds.\n\nBeyond autoscaling, Cast AI automates MIG configuration and time-slicing deployment (enabled in Cast AI Settings → Workload Management → GPU Optimization), so platform engineers do not manually tune MIG partition profiles per node. The [Kubernetes Scheduler for GPU-Heavy Apps with Node Templates](https://cast.ai/blog/kubernetes-scheduler-for-gpu-heavy-apps-with-node-templates/) post covers how node templates constrain autoscaling to specific GPU instance types that map directly to your Kueue ResourceFlavors, keeping the ResourceFlavor-to-node-group mapping accurate as the cluster scales.\n\n## Conclusion\n\nGPU utilization averaging 5% is not a hardware problem. It is a queue problem and a device description problem. Kueue solves the first: jobs wait in an observable, orderable, preemptable queue rather than piling up as zombie Pending pods. DRA solves the second: workloads describe exactly what they need, and the scheduler places them on the right device slice rather than the first whole card that fits. Start with a single ClusterQueue, one ResourceFlavor per GPU node group, and a LocalQueue per team. Add DRA device classes for MIG slices once your NVIDIA GPU Operator version supports the DRA driver. Measure queue depth and admission latency via Kueue’s Prometheus metrics. Then scale the model across additional teams.\n\nA cluster queuing 40 GPU jobs with 12 teams no longer needs an all-hands Slack thread every Monday morning. Kueue enforces the fairness policy; DRA places the right job on the right device; Cast AI provisions the nodes and scales down between runs. See [Cast AI GPU Optimization](https://cast.ai/gpu-optimization/) for how automated GPU optimization integrates with this scheduling stack.\n\n## Frequently Asked Questions\n\n### **What is Kueue?**\n\nKueue is a Kubernetes-native job queuing controller (CNCF Incubating, kubernetes-sigs). It acts as an admission controller above kube-scheduler: jobs are created in a Suspended state, held in a queue, and admitted only when quota and node capacity are both available. Kueue does not replace kube-scheduler. Instead, it gates access to it, making GPU capacity allocation observable, fair, and preemptable across teams.\n\n### **What is Dynamic Resource Allocation in Kubernetes?**\n\nDynamic Resource Allocation (DRA) is a Kubernetes API that [graduated to GA in Kubernetes 1.34](https://kubernetes.io/blog/2025/09/01/kubernetes-v1-34-dra-updates/) (KEP-3063) and is enabled by default from 1.34 onward. DRA lets workloads describe hardware requirements as structured attributes via ResourceClaim objects—for example, “a GPU with at least 40GB VRAM and Ampere architecture”—rather than requesting opaque integer counts. The scheduler matches claims to ResourceSlices published by device drivers, enabling precise placement and MIG slice selection.\n\n### **How do I queue GPU jobs in Kubernetes?**\n\nKueue GPU scheduling works by adding the label `kueue.x-k8s.io/queue-name: <your-LocalQueue-name>` to your batch/v1 Job, JobSet, PyTorchJob, or RayJob manifest. Create a LocalQueue in your namespace that points to a ClusterQueue with GPU nominalQuota configured. Kueue holds your job until GPU quota is available and flips spec.suspend=false to release it to kube-scheduler for pod placement. No changes to container images or application code are required.\n\n### **Kueue vs Volcano: which should I use?**\n\nUse Kueue for multi-tenant GPU batch, LLM fine-tuning queues, and agent evaluation sweeps. Kueue is CRD-based, kubectl-native, and runs on top of kube-scheduler with no second scheduler binary. Use Volcano when you need MPI-style gang scheduling at the pod-placement level, or when your teams operate HPC workloads that expect Volcano’s PodGroup primitives. Kueue 1.1+ plus JobSet closes most of the practical gap for AI workloads.\n\n### **Does Kueue support gang scheduling?**\n\nYes, at the admission level. Kueue admits a Workload atomically: all pod sets are admitted together, or none are. Kueue 1.1+ provides first-class support for JobSet, giving pod-level gang coordination with shared failure, restart, and success handling on top of the default scheduler. For strict MPI gang scheduling at the pod-placement level, pair Kueue with Volcano.\n\n### **How do I set GPU quotas per team in Kueue?**\n\nCreate one ClusterQueue per team with the desired nominalQuota for your GPU resource, for example `nvidia.com/gpu: 8`, or a DRA DeviceClass name mapped via deviceClassMappings. Group ClusterQueues into a cohort to enable borrowing. Set borrowingLimit to cap how much a team can burst into idle capacity, and lendingLimit to control how much they share. Enable `reclaimWithinCohort: Any` to ensure each team always gets their nominal GPU allocation back when they submit new work, regardless of who is currently borrowing.", "url": "https://wpnews.pro/news/gpu-job-queueing-with-kueue-and-dra-scheduling-ai-workloads-without-idle", "canonical_source": "https://cast.ai/blog/kueue-gpu-scheduling-cost/", "published_at": "2026-09-25 15:12:20+00:00", "updated_at": "2026-10-02 09:36:18.272732+00:00", "lang": "en", "topics": ["ai-infrastructure", "mlops", "ai-chips"], "entities": ["Cast AI", "Kueue", "Kubernetes", "Dynamic Resource Allocation", "JobSet", "LeaderWorkerSet", "H100", "H200"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/gpu-job-queueing-with-kueue-and-dra-scheduling-ai-workloads-without-idle", "markdown": "https://wpnews.pro/news/gpu-job-queueing-with-kueue-and-dra-scheduling-ai-workloads-without-idle.md", "text": "https://wpnews.pro/news/gpu-job-queueing-with-kueue-and-dra-scheduling-ai-workloads-without-idle.txt", "jsonld": "https://wpnews.pro/news/gpu-job-queueing-with-kueue-and-dra-scheduling-ai-workloads-without-idle.jsonld"}}