Gemma 4 on Amazon SageMaker: The NVIDIA T4 Decodes at 0.8x of the L4 With the Same Answers A developer deployed Gemma 4 (E2B, E4B, 12B and 26B A4B) to Amazon SageMaker on a single NVIDIA T4 (ml.g4dn.xlarge), measuring decode throughput at 0.77x to 0.82x of an L4 while producing identical answers, with 12B the largest build that fits in 16 GB. The work required a Turing attention patch to vLLM 0.30.0, since the Triton attention kernel requests 98,304 bytes of shared memory per block against the T4's 65,536-byte limit, and includes a 26B A4B W4A16 repack that keeps Google's QAT weights on their 4-bit grid (0 of 11,534,336 layer-0 query-projection weights change level, versus 30.3% for an AWQ build). A companion MCP server exposes tools including check_quotas for reading per-region SageMaker endpoint quotas. This article deploys Gemma 4 to the smallest GPU Amazon SageMaker offers, an NVIDIA T4, and measures it against the L4 from the earlier parts. A suite of Python MCP tools is built to simplify management of the vLLM hosted deployment. https://github.com/xbill9/sagemaker-gemma https://github.com/xbill9/sagemaker-gemma | Models | Gemma 4 E2B, E4B, 12B and 26B A4B, 4-bit weights and 4-bit embeddings | | Hardware | SageMaker ml.g4dn.xlarge , 1x NVIDIA T4, 16 GB; compared with ml.g6.xlarge , 1x L4 | | Region | us-east-2 | | Software | AWS vLLM SageMaker container 0.30.0, with a Turing attention patch | | Result | The T4 decodes at 0.77x to 0.82x of the L4 and gives identical answers; 12B is the largest build that fits | This is part four of a series. Part one deploys Gemma 4 to a SageMaker endpoint with the aws CLI and an MCP server: https://dev.to/aws-builders/gemma-4-on-an-amazon-sagemaker-endpoint-aws-cli-nvidia-l4-and-an-mcp-server-2c9d https://dev.to/aws-builders/gemma-4-on-an-amazon-sagemaker-endpoint-aws-cli-nvidia-l4-and-an-mcp-server-2c9d Part two measures Google's QAT checkpoint against the full-size bf16 release: https://dev.to/aws-builders/gemma-4-on-amazon-sagemaker-qat-weights-decode-205x-faster-than-bf16-on-one-l4-318m https://dev.to/aws-builders/gemma-4-on-amazon-sagemaker-qat-weights-decode-205x-faster-than-bf16-on-one-l4-318m Part three repacks the QAT weights with 4-bit embeddings: https://dev.to/aws-builders/gemma-4-on-amazon-sagemaker-4-bit-embeddings-decode-up-to-139x-faster-on-one-l4-36mf https://dev.to/aws-builders/gemma-4-on-amazon-sagemaker-4-bit-embeddings-decode-up-to-139x-faster-on-one-l4-36mf This article takes part three's builds down one GPU generation. SageMaker JumpStart lists Gemma 4 from ml.g6e.xlarge , one NVIDIA L40S, for E2B, and 12B only on ml.g6e.16xlarge ; none of its Gemma 4 entries lists a T4. Gemma 4 on Turing GPUs is an open vLLM issue, 38918, "Gemma4 on Turing GPUs SM 7.5 : all attention backends hit shared memory limits", reported again on vLLM 0.29.0 in September. The builds served here, with 4-bit embeddings, are the author's repacks of Google's QAT weights on Hugging Face. The 26B A4B build is the first W4A16 checkpoint of that model to keep Google's QAT weights on their 4-bit grid. In layer 0's query projection, 0 of its 11,534,336 weights change level, where an AWQ build of the same QAT export, published in June, moves 30.3% of them by a full level or more. grid check.py reads that one tensor from each checkpoint over HTTP and compares them. make test passing aws login session in an account with SageMaker endpoint quota for ml.g4dn.xlarge The project's MCP server has a check quotas tool that reads the account's endpoint quota for each instance type in every US region: | Instance | us-east-1 | us-east-2 | us-west-1 | us-west-2 | | ml.g4dn.xlarge | 2 | 2 | 2 | 2 | | ml.g5g.xlarge | - | - | - | - | | ml.g5.xlarge | 2 | 2 | - | 2 | | ml.g6.xlarge | 1 | 1 | - | 1 | | ml.inf2.xlarge | 0 | 2 | - | 0 | A dash means SageMaker has no quota entry for the type. ml.g4dn.xlarge , one T4 with 16 GB, is the smallest GPU on offer. The Graviton T4G g5g exists on EC2 only. Inferentia2 is listed, but it runs the Neuron SDK and needs a different container and a compiled model. Gemma 4's full-attention layers are 512 wide. vLLM serves them with its Triton attention kernel, whose tile asks for 98,304 bytes of shared memory per block. A T4 is a Turing GPU and allows 65,536, so the engine stops at start-up: triton.runtime.errors.OutOfResources: shared memory, Required: 98304, Hardware limit: 65536 vLLM 0.30.0, the version in the SageMaker container, has no fix for it. The author's Compute Engine T4 serves Gemma 4 with a script that halves the tiles on pre-Ampere GPUs, and the SageMaker image gets the same script, in a Dockerfile built FROM the stock container: ARG BASE IMAGE FROM ${BASE IMAGE} COPY patch triton turing.py /opt/turing/patch triton turing.py RUN set -eu; \ target="$ python3 -c 'import importlib.util, os; print os.path.join importlib.util.find spec "vllm" .submodule search locations 0 , "v1/attention/ops/triton unified attention.py" ' "; \ python3 /opt/turing/patch triton turing.py "$target"; \ python3 /opt/turing/patch triton turing.py --check "$target"; \ python3 -c 'import torch, sys; a = torch. C. cuda getArchFlags ; print "torch arch:", a ; sys.exit 0 if "sm 75" in a else 1 ' The last line refuses the build if PyTorch in the image has no Turing kernels. CodeBuild builds the image and pushes it to a private ECR repository, so the 8 GB base image never touches the local disk: 8 1.092 patch triton turing: patched /usr/local/lib/python3.12/dist-packages/vllm/v1/attention/ops/triton unified attention.py 8 1.092 smem budget : 60000 B of Turing's 65536 hard limit 8 5.652 torch arch: sm 75 sm 80 sm 86 sm 90 sm 100 sm 120 The patch is a no-op on an L4 or newer, and the image keeps the stock entrypoint, so every SM VLLM setting works as before. The Dockerfile and buildspec.yml are in the repository, in the turing directory. A SageMaker production variant runs on one of several host images, each with its own NVIDIA driver, chosen by InferenceAmiVersion : al2-ami-sagemaker-inference-gpu-2 NVIDIA driver 535, CUDA 12.2 al2-ami-sagemaker-inference-gpu-3-1 NVIDIA driver 550, CUDA 12.4 al2023-ami-sagemaker-inference-gpu-4-1 NVIDIA driver 580, CUDA 13.0 The vLLM 0.30.0 container is built on CUDA 13: NVIDIA REQUIRE CUDA=cuda =13.0 ... CUDA VERSION=13.0.2 sm.py passes the host image through from .env : INFERENCE AMI VERSION=al2023-ami-sagemaker-inference-gpu-4-1 Without InferenceAmiVersion , an ml.g4dn.xlarge endpoint with this image ends about six minutes after Creating with: CannotStartContainerError. Please ensure the model container for variant AllTraffic starts correctly when invoked with 'docker run