cd /news/ai-infrastructure/gemma-4-on-amazon-sagemaker-or-a-vm-… · home › topics › ai-infrastructure › article
[ARTICLE · art-142709] src=dev.to ↗ pub= topic=ai-infrastructure verified=true sentiment=· neutral

Gemma 4 on Amazon SageMaker or a VM? The Same Model Server at 1.40x the Price

A developer benchmarked the same Gemma 4 E2B build served from an Amazon SageMaker endpoint and from a plain EC2 instance with an identical NVIDIA T4 or L4 GPU, finding decode speed and answers match within 2% while EC2 costs 0.71x per hour and its on-instance client pays 0.006 s per call versus 0.56 s through the aws CLI. The EC2 rig, built from a cloud-init script that applies the same Turing patch as the SageMaker path, began serving 9.7 minutes after launch with no inbound security-group rules, using Systems Manager as the only access path. A suite of Python MCP tools was built to simplify management of the vLLM-hosted deployment.

by read7 min views3 publishedSep 30, 2026

This article serves the same Gemma 4 build from an Amazon SageMaker endpoint and from a plain EC2 instance with the same GPU, on an NVIDIA T4 and an NVIDIA L4, and prices both against Compute Engine and Cloud Run. A suite of Python MCP tools is built to simplify management of the vLLM hosted deployment.

https://github.com/xbill9/sagemaker-gemma

| Model | Gemma 4 E2B, 4-bit weights (4-bit embeddings on the T4) | | Hardware | 1x NVIDIA T4 ( ml.g4dn.xlarge /g4dn.xlarge ) and 1x NVIDIA L4 (ml.g6.xlarge /g6.xlarge ) | | Region | us-east-2 | | Software | vLLM 0.30.0 on both sides | | Result | Decode and answers match within 2%. EC2 costs 0.71x per hour, and its client, on the instance, pays0.006 s per call against0.56 s through the aws CLI |

This is part five of a series. Part one deploys Gemma 4 to a SageMaker endpoint with the aws CLI and an MCP server: https://dev.to/aws-builders/gemma-4-on-an-amazon-sagemaker-endpoint-aws-cli-nvidia-l4-and-an-mcp-server-2c9d

Part two measures Google's QAT checkpoint against the full-size bf16 release: https://dev.to/aws-builders/gemma-4-on-amazon-sagemaker-qat-weights-decode-205x-faster-than-bf16-on-one-l4-318m

Part three repacks the QAT weights with 4-bit embeddings: https://dev.to/aws-builders/gemma-4-on-amazon-sagemaker-4-bit-embeddings-decode-up-to-139x-faster-on-one-l4-36mf

Part four serves those builds on SageMaker's smallest GPU, the T4: https://dev.to/aws-builders/gemma-4-on-amazon-sagemaker-the-nvidia-t4-decodes-at-08x-of-the-l4-with-the-same-answers-hm7

This article asks what the managed endpoint adds over the same GPU without it.

make test passingaws login session with SageMaker endpoint quota and EC2 on-demand G-family vCPU quota in us-east-2 Two ways to serve one model on one GPU:

SageMaker endpoint EC2 instance
Provisioning create-endpoint run-instances with cloud-init
vLLM AWS SageMaker container 0.30.0 vllm/vllm-openai:v0.30.0
Turing patch on the T4 derived image, built by CodeBuild derived image, built on the instance
Access IAM-signed invoke-endpoint Systems Manager only, no inbound rules
Client aws CLI on a workstation compare.py on the instance, to localhost

The model, context length (8,192), memory setting (0.90), data type and measurement script are the same on both sides. The client location differs, and the per-call figures below include it.

The EC2 side is a rig in the author's gemma4-dev tree, gpu-vllm-g4dn-2b-w4a16, whose settings match the SageMaker endpoint:

INSTANCE_TYPE=g4dn.xlarge
VLLM_IMAGE=vllm/vllm-openai:v0.30.0
VLLM_PATCHED_IMAGE=vllm-openai:v0.30.0-sm75-patched
DTYPE=float16
MAX_MODEL_LEN=8192
GPU_MEMORY_UTILIZATION=0.90
MODEL=xbill9/gemma-4-E2B-it-qat-q4_0-w4a16-ct-text-emb4

Cloud-init pulls the stock image, applies the same Turing patch as part four and builds the derived tag on the instance, in seconds:

[stage] image-pull-done +170s
[stage] patch-applied +191s
[stage] image-build-done +194s
[stage] patch-verified-in-image +209s
[stage] serving-started +211s

The instance served 9.7 minutes after launch. Its security group has no inbound rules; Systems Manager is the only way in.

ec2_measure.py copies the project's compare.py onto the instance over Systems Manager and runs it against localhost:8000. It measures decode speed, 1 to 16 parallel requests and 40 questions at temperature 0, the same measurement every SageMaker endpoint in the series got. The instance is terminated in a finally block, and a watchdog terminates it if the driver dies.

compare.py combine sets the two runs side by side:

                            gemma-4-e2b-emb4-t4  gemma-4-e2b-emb4-g4dn  ratio
weights_gib                           2.86                2.86  1.0
kv_cache_tokens                     660033              660108  1.0
decode_tokens_per_second             108.5               111.2  1.02
load_c1_tokens_per_second             91.1              113.85  1.25
load_c4_tokens_per_second           308.25               394.9  1.28
load_c16_tokens_per_second          787.05              1005.1  1.28
quality_correct                         36                  36
identical answers: 40/40

The GPU does the same work on both: the same memory allocation, decode within 2% and all 40 answers byte-identical.

The same comparison on the L4, with Google's E2B QAT checkpoint and the stock vLLM 0.30.0 image:

Measure SageMaker ml.g6.xlarge EC2 g6.xlarge EC2 / SageMaker
Decode, one request (tok/s) 105.1 105.1 1.00
1 request, 256 tokens (tok/s) 85.35 104.85 1.23
16 parallel (tok/s) 1077.25 1485.9 1.38
Per-call fixed cost (s) 0.562 0.008 –
Identical answers 40 of 40

Two GPUs, two checkpoints, one result: the model server runs at the same speed on SageMaker and on EC2.

compare.py fits each request's wall time as a fixed cost plus a per-token rate. The per-token rate is the decode speed, and it matches. The fixed cost is 0.56 s per call through SageMaker and under 0.01 s on the instance.

That 0.56 s covers everything between the client and vLLM: starting the aws CLI, signing the request, the network round trip from the workstation and the SageMaker front end. The EC2 figure has none of the first three, because the client runs beside vLLM. A client off the instance would pay its own network cost, and a long-lived SDK client would skip the CLI start-up; neither was measured here. The fixed cost per call is what lowers SageMaker's rate at 1 to 16 parallel requests, by 1.23x to 1.38x.

On-demand list prices, from the AWS Price List API and the Google Cloud billing catalog on 2026-09-30:

Option T4 ($/h) L4 ($/h)
SageMaker endpoint 0.736 1.1267
EC2 0.526 0.8048
Compute Engine VM 0.5241 0.7045
Cloud Run, per running hour – 1.4209

SageMaker costs 1.40x EC2 per hour on both GPUs. The Compute Engine T4 is an n1-standard-2 in us-west2; the L4 is a g2-standard-4 in us-east4. Cloud Run is one L4 with 8 vCPU and 32 GiB in us-east4, instance-based billing, no zonal redundancy.

Per million output tokens at 16 parallel requests, each platform at its own measured rate:

GPU SageMaker EC2 EC2 / SageMaker
T4 $0.260 $0.145 0.56
L4 $0.291 $0.150 0.52

At equal throughput, the price alone makes EC2 0.71x. The rest comes from the faster call path on the instance.

Cloud Run's L4 costs 1.67x a Compute Engine g2-standard-8 of the same size for every hour it runs, and with --min-instances=0 it runs only while there is traffic. By arithmetic on the list prices, it is the cheaper of the two when the instance is up less than 60% of the day, about 14.4 hours. Each start after a quiet period loads the model again, which takes minutes for Gemma 4, so scale-to-zero suits a model that is called in bursts. Throughput on Compute Engine and Cloud Run was not measured here.

SM_VLLM_MODEL, SM_VLLM_MAX_MODEL_LEN and the rest become vLLM flags, on a container AWS maintains. On an L4 or newer GPU there is nothing to build.invoke-endpoint. InstancePools tries the next instance type of the same GPU when the first has no capacity.InService on the T4, against 9.7 minutes to serving on EC2 with the image built on the instance.

Stage Where Why
🥇 Finding a working vLLM configuration SageMaker endpoint Maintained container, flags as settings, logs and capacity fallback
🥇 Serving a small model after that EC2 or Compute Engine Same model server at 0.71x the hourly price
🥈 Bursty traffic on Google Cloud Cloud Run Scales to zero
🥇 Traffic that needs scaling and safe updates SageMaker endpoint The features the 1.40x pays for

The endpoint is a quick way to find a vLLM configuration that works, with the logs and fallback to debug it. Once the settings are known, a model that fits one GPU serves at the same speed from a VM for 0.71x the price.

The EC2 instance was terminated when the measurement finished:

2026-09-30T16:21:21Z measured
2026-09-30T16:28:11Z i-08bf5db2ef4c05ffb terminated

The SageMaker endpoints were deleted after their measurements, as in part four.

The goal of this article was to find what a SageMaker endpoint adds over the same GPU without it. The key to the solution was serving the same build with the same vLLM version both ways and measuring both with one script. The results were:

Scope: one account, us-east-2, one deployment of each on 2026-09-29 (L4) and 2026-09-30 (T4), vLLM 0.30.0 throughout. The SageMaker client was the aws CLI on a workstation, and the EC2 client ran on the instance, so the per-call figures include the client's location; a remote client for EC2 and an SDK client for SageMaker were not measured. The SageMaker T4 ran the AWS container with a derived patch and the EC2 T4 the public vLLM image with the same patch. Prices are on-demand list prices on 2026-09-30, with no Savings Plans, Spot, sustained-use or committed-use discounts, which differ by platform.

The strategy for using MCP for SageMaker deployment and benchmarking was validated with an incremental step by step approach.

── more in #ai-infrastructure 4 stories · sorted by recency
── more on @amazon sagemaker 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/gemma-4-on-amazon-sa…] indexed:0 read:7min 2026-09-30 · —