Let Karpenter provision and bin-pack spot GPU nodes under a vLLM server, then delete them when traffic drops.
What you'll build #
A Kubernetes-hosted inference stack on Amazon EKS where Karpenter provisions spot GPU nodes on demand, bin-packs a vLLM model server onto the cheapest available capacity, and deletes the nodes when traffic drops — no static GPU node group burning money at 3 a.m.
Prerequisites #
Verified against Karpenter v1.14.0, Kubernetes 1.36 on EKS, NVIDIA device plugin v0.19.3, and vLLM v0.25.1 in July 2026. You need:
- An
EKScluster (Kubernetes 1.33+) with Karpenter v1.14.0 already installed per theofficial getting-started guide— that guide sets up the
KarpenterNodeRole-<cluster>
IAM role, the spot-interruption SQS queue (settings.interruptionQueue
), andkarpenter.sh/discovery
tags on your subnets and security groups. This tutorial assumes all of that exists. kubectl,Helm3.14+, andAWS CLIv2 authenticated against the cluster's account.- A region with G5/G6 capacity (e.g.
us-east-1
,us-west-2
). - An EC2 spot quota above zero for GPU instances — check
All G and VT Spot Instance Requests(quota codeL-3819A6DF
). Fresh accounts often have 0, which blocks everything below:
aws service-quotas get-service-quota \
--service-code ec2 --quota-code L-3819A6DF \
--query 'Quota.Value'
If it returns 0.0
, request an increase (8+ vCPUs) before starting. Export your cluster name for the steps below:
export CLUSTER_NAME=<your-cluster>
1. Enable spot in the account #
If this account has never launched a spot instance, EC2 needs its service-linked role:
aws iam create-service-linked-role --aws-service-name spot.amazonaws.com
An InvalidInput
error saying the role name "has been taken" means it already exists — safe to ignore.
2. Create a GPU EC2NodeClass #
The EC2NodeClass tells Karpenter how to build GPU nodes. The al2023@latest
AMI alias is doing real work here: Karpenter resolves it to the EKS-optimized AL2023 AMI variants — including the NVIDIA accelerated one, which ships the GPU drivers and container toolkit — and picks the right variant per instance type automatically. The 100 GiB root volume matters because the vLLM image alone is over 10 GiB; the AL2023 default of 20 GiB gets you disk-pressure evictions.
cat <<EOF | kubectl apply -f -
apiVersion: karpenter.k8s.aws/v1
kind: EC2NodeClass
metadata:
name: gpu-inference
spec:
role: "KarpenterNodeRole-${CLUSTER_NAME}"
amiSelectorTerms:
- alias: al2023@latest
subnetSelectorTerms:
- tags:
karpenter.sh/discovery: "${CLUSTER_NAME}"
securityGroupSelectorTerms:
- tags:
karpenter.sh/discovery: "${CLUSTER_NAME}"
blockDeviceMappings:
- deviceName: /dev/xvda
ebs:
volumeSize: 100Gi
volumeType: gp3
EOF
3. Create the spot GPU NodePool #
The NodePool defines what Karpenter may launch. Two families (g6
= NVIDIA L4, g5
= A10G) across all sizes gives EC2 Fleet many spot pools to choose from — diversification is what keeps spot interruption rates low. Allowing both capacity types means Karpenter uses spot when it's available and falls back to on-demand only when it isn't; Fleet picks the actual pool with the price-capacity-optimized strategy (lowest price, lowest interruption risk). The taint keeps non-GPU workloads off these expensive nodes, and the limits
block caps the blast radius at 8 GPUs no matter what gets deployed.
cat <<EOF | kubectl apply -f -
apiVersion: karpenter.sh/v1
kind: NodePool
metadata:
name: gpu-inference
spec:
template:
spec:
taints:
- key: nvidia.com/gpu
value: "true"
effect: NoSchedule
requirements:
- key: kubernetes.io/arch
operator: In
values: ["amd64"]
- key: kubernetes.io/os
operator: In
values: ["linux"]
- key: karpenter.sh/capacity-type
operator: In
values: ["spot", "on-demand"]
- key: karpenter.k8s.aws/instance-family
operator: In
values: ["g6", "g5"]
nodeClassRef:
group: karpenter.k8s.aws
kind: EC2NodeClass
name: gpu-inference
limits:
nvidia.com/gpu: "8"
disruption:
consolidationPolicy: WhenEmptyOrUnderutilized
consolidateAfter: 5m
EOF
consolidateAfter: 5m
gives you five minutes of headroom before Karpenter starts tearing down underutilized nodes — long enough to survive a rolling restart without churning instances.
4. Install the NVIDIA device plugin #
The accelerated AMI has drivers, but Kubernetes only learns about the GPUs from the NVIDIA device plugin. Karpenter also won't consider a GPU node fully initialized until the nvidia.com/gpu
resource registers, so this DaemonSet is load-bearing. Its default values already tolerate the nvidia.com/gpu:NoSchedule
taint:
helm repo add nvdp https://nvidia.github.io/k8s-device-plugin
helm repo update
helm upgrade -i nvdp nvdp/nvidia-device-plugin \
--namespace nvidia-device-plugin --create-namespace \
--version 0.19.3
5. Deploy the model server #
A vLLM deployment serving Qwen2.5-0.5B-Instruct — small enough to download in a minute and fit any single L4/A10G. The nvidia.com/gpu: "1"
request is the trigger: the pod goes Pending, Karpenter sees an unschedulable GPU pod, and launches a node for it. The in-memory /dev/shm
volume avoids vLLM's shared-memory warnings; the CPU/memory requests are sized so one replica bin-packs cleanly onto an xlarge
.
cat <<EOF | kubectl apply -f -
apiVersion: apps/v1
kind: Deployment
metadata:
name: vllm
spec:
replicas: 1
selector:
matchLabels:
app: vllm
template:
metadata:
labels:
app: vllm
spec:
tolerations:
- key: nvidia.com/gpu
operator: Exists
effect: NoSchedule
containers:
- name: vllm
image: vllm/vllm-openai:v0.25.1
args:
- --model=Qwen/Qwen2.5-0.5B-Instruct
- --max-model-len=4096
ports:
- containerPort: 8000
resources:
requests:
cpu: "3"
memory: 10Gi
nvidia.com/gpu: "1"
limits:
nvidia.com/gpu: "1"
readinessProbe:
httpGet:
path: /health
port: 8000
initialDelaySeconds: 60
periodSeconds: 10
volumeMounts:
- name: shm
mountPath: /dev/shm
volumes:
- name: shm
emptyDir:
medium: Memory
sizeLimit: 2Gi
EOF
Verify it works #
Watch Karpenter react — a NodeClaim should appear within seconds and go Ready in about two minutes:
kubectl get nodeclaims -w
NAME TYPE CAPACITY ZONE NODE READY AGE
gpu-inference-8xkzt g6.xlarge spot us-east-1b ip-192-168-73-21.ec2.internal True 2m4s
CAPACITY: spot
is the money column. Once the pod is Ready (first start takes a few minutes: image pull plus model download), hit the OpenAI-compatible API:
kubectl port-forward deploy/vllm 8000:8000 &
curl -s localhost:8000/v1/models
Expected output (trimmed):
{"object":"list","data":[{"id":"Qwen/Qwen2.5-0.5B-Instruct","object":"model","owned_by":"vllm", ...}]}
Now watch the bin-packing. Scale up and Karpenter batches the pending pods, then asks EC2 Fleet for the cheapest combination that fits them — several single-GPU spot nodes or one multi-GPU node, whichever prices lower:
kubectl scale deploy/vllm --replicas=3
kubectl get nodeclaims
Scale back to 1 and, five minutes later (consolidateAfter
), watch Karpenter consolidate the now-empty nodes out of existence. That's the whole cost story: GPU capacity exists only while pods need it, and it's spot-priced — AWS quotes up to 90% off on-demand, and G-family discounts routinely land north of 50%.
Troubleshooting #
Karpenter logs show MaxSpotInstanceCountExceeded during instance launch — your spot vCPU quota for G instances is 0 or too low. Request an increase on
L-3819A6DF
(All G and VT Spot Instance Requests); until it lands, Karpenter will retry and may fall back to on-demand.Pod event: incompatible with nodepool "gpu-inference" ... no instance type satisfied resources — the pod's requests can't fit any instance the NodePool allows, or the NodePool hit its
limits
. Check that requests fit an xlarge
(4 vCPU/16 GiB) and that existing NodeClaims haven't consumed the 8-GPU cap.Node is Ready but the pod stays Pending with Insufficient nvidia.com/gpu — the device plugin isn't running on the node.
kubectl get pods -n nvidia-device-plugin -o wide
should show one pod per GPU node; if you overrode chart values, make sure the DaemonSet still tolerates nvidia.com/gpu:NoSchedule
.vLLM CrashLoopBackOff with torch.OutOfMemoryError: CUDA out of memory — the model plus KV cache doesn't fit the GPU. Lower
--max-model-len
, pass --gpu-memory-utilization 0.85
, or constrain the NodePool to instance types with more VRAM.## Next steps
Pin the AMI alias to a specific version (al2023@v2026...
) in production — @latest
triggers drift replacement of GPU nodes on every AMI release, which can churn long-running inference pods. Add a second, weighted NodePool with on-demand
only to guarantee a baseline while spot handles overflow. Karpenter only scales nodes; pair it with an HPA or KEDA scaling on queue depth or request latency so replicas track traffic. From there, dig into disruption budgets to bound how many nodes consolidate at once, and NVIDIA time-slicing or MIG in the device plugin to pack multiple small models onto one GPU.
Sources & further reading #
Getting Started with Karpenter (v1.14)— karpenter.sh - Karpenter NodeClasses— karpenter.sh - Karpenter Scheduling Concepts— karpenter.sh - NVIDIA device plugin for Kubernetes— github.com - vLLM Releases— github.com - Using Amazon EC2 Spot Instances with Karpenter— aws.amazon.com
Emeka Okafor· Security Editor
Emeka has spent over a decade tracking threat actors, vulnerability disclosures, and the evolving landscape of application security, bringing a sharp continent-spanning perspective to his reporting. He's known for translating dense CVE advisories into clear, actionable context that developers and security teams alike actually read.
Discussion 0 #
No comments yet
Be the first to weigh in.