Gemma 4 Inference on AWS: Bedrock, SageMaker, GPUs, Inferentia and Trainium Behind One Strands Agent A developer benchmarked six AWS LLM inference backends — Amazon Bedrock, a SageMaker real-time endpoint, vLLM on two EC2 GPU families, and a hand-ported Gemma 4 build on Inferentia2 and Trainium — all driven by a single Strands agent using the same prompts and grading scripts. The survey documents setup, failure modes and pricing per path, including capacity errors on inf2.xlarge in us-east-2a and Trainium spot quota requests that opened support cases in four US regions without automatic approval. The g6 vLLM launch took 14.0 minutes to become healthy, and the author notes Bedrock's inference profile id (us.amazon.nova-micro-v1:0) is required for on-demand throughput. This article provides a step by step survey of LLM inference on AWS: Amazon Bedrock, a SageMaker real-time endpoint, vLLM on two EC2 GPU families, and a hand-ported Gemma 4 on AWS Inferentia2 and Trainium. A Strands agent drives all six backends, and every number below comes from two complete runs on the same day with the same prompts. https://github.com/xbill9/aws-inference-strands https://github.com/xbill9/aws-inference-strands AWS offers a model as an API call, as a managed endpoint, as a GPU you rent by the hour, and as two families of its own accelerator chips. Each path has its own setup, its own failure modes and its own price, and most comparisons cover one or two of them. This survey puts all of them behind one agent. The same Strands code talks to Bedrock, to SageMaker, to vLLM on EC2, and to a Neuron server, and the same scripts grade the answers and time the requests. | Backend | What serves the model | Model | |---|---|---| | Amazon Bedrock | managed API | us.amazon.nova-micro-v1:0 | | EC2 g6.xlarge NVIDIA L4 | vLLM 0.30 | xbill9/gemma-4-E2B-it-qat-q4 0-w4a16-ct-text-emb4 | | SageMaker ml.g6.xlarge | AWS vLLM 0.30 container | the same repack | | EC2 g5g.2xlarge Graviton2 + T4G | vLLM 0.27.2rc0, patched for sm 75 | google/gemma-4-E2B-it , fp16 | | EC2 inf2.xlarge Inferentia2 | hand-ported Gemma 4, xbill9/gemma4-optb:slim | Gemma 4 E2B, compiled for Neuron | | EC2 trn1.2xlarge Trainium | the same image, unchanged | the same build | The g6 and SageMaker rows run the same weights on the same GPU, one as a VM and one as a managed endpoint. The inf2 and trn1 rows run the same compiled image on two different AWS chips. AmazonSSMManagedInstanceCore , and a security group with no inbound rules. pip install 'strands-agents openai,sagemaker ' boto3 . Quota is permission to ask. Capacity is whether AWS has the machine. Both decide whether a run happens today. us-east-2 Running On-Demand Trn instances 8.0 us-east-2 All Trn Spot Instance Requests 0.0 us-east-2 Trn spot request: CASE OPENED 8.0 2026-10-04T21:36:11.599000-04:00 Eight vCPUs of on-demand Trainium buys exactly one trn1.2xlarge . Spot starts at zero. A request for eight vCPUs of spot in each of the four US regions opened a support case in each one, and none was approved automatically. The instance type itself is offered in one zone of the three regions checked: us-east-2c . An inf2.xlarge launch in us-east-2a , with 80 vCPUs of Inferentia quota free, returned: An error occurred InsufficientInstanceCapacity when calling the RunInstances operation: We currently do not have sufficient inf2.xlarge capacity in the Availability Zone you requested us-east-2a . The same launch in us-east-2b succeeded. Bedrock is the reference point: no instance, no endpoint, an IAM permission and a model id. The inference profile id us.amazon.nova-micro-v1:0 is the one to use; the bare model id returns a validation error for on-demand throughput. python from strands.models import BedrockModel model = BedrockModel model id="us.amazon.nova-micro-v1:0", region name="us-east-1", temperature=0.0 The g6 serves the 4-bit embedding repack of Gemma 4 E2B with vLLM's OpenAI server. Tool calls need two flags; without them the agent's tools are ignored. docker run -d --gpus all --ipc=host -p 8000:8000 vllm/vllm-openai:v0.30.0 \ --model xbill9/gemma-4-E2B-it-qat-q4 0-w4a16-ct-text-emb4 \ --max-model-len 8192 --gpu-memory-utilization 0.90 \ --enable-auto-tool-choice --tool-call-parser gemma4 g6 vLLM 0.30, emb4 : launch 2026-10-05T02:02:48+00:00 healthy 2026-10-05T02:16:46+00:00 - 14.0 min Strands reaches it with the OpenAI provider: python from strands.models.openai import OpenAIModel model = OpenAIModel client args={"base url": "http://localhost:8001/v1", "api key": "unused"}, model id="xbill9/gemma-4-E2B-it-qat-q4 0-w4a16-ct-text-emb4" The g5g pairs a Graviton2 Arm host with an NVIDIA T4G: Arm and CUDA in an instance that costs $0.556 an hour. vLLM's arm64 image ships without the T4G's sm 75 kernels, and Gemma 4's 512-wide attention heads need more shared memory than Turing has, so this row runs a vLLM 0.27.2rc0 built from source for sm 75 with a shared-memory patch, from a prebuilt AMI. The same two tool flags go on its command line. EngineCore pid=2128 INFO 10-05 02:55:22 default loader.py:430 Loading weights took 520.97 seconds APIServer pid=1651 INFO 10-05 02:58:21 parser manager.py:37 "auto" tool choice has been enabled. The 521 seconds of weight loading come from the first read of a fresh volume restored from the AMI snapshot. The SageMaker row uses the AWS vLLM container on ml.g6.xlarge , the same GPU as the EC2 g6, with vLLM flags passed as SM VLLM environment variables: { "SM VLLM ENABLE AUTO TOOL CHOICE": "true", "SM VLLM MODEL": "xbill9/gemma-4-E2B-it-qat-q4 0-w4a16-ct-text-emb4", "SM VLLM TOOL CALL PARSER": "gemma4" } SageMaker emb4 endpoint: make deploy returned at epoch 1791166343, InService at epoch 1791166834 aws sagemaker wait - 8.2 min Strands reaches it with SageMakerAIModel , which signs invoke endpoint calls instead of opening a URL: python from strands.models.sagemaker import SageMakerAIModel model = SageMakerAIModel endpoint config={"endpoint name": "gemma-4-e2b-emb4-strands", "region name": "us-east-2"}, payload config={"max tokens": 512, "stream": False, "temperature": 0.0} No vLLM release serves Gemma 4 on Neuron, so this row runs a hand-ported Gemma 4: a torch neuronx graph compiled ahead of time and baked into a container with an OpenAI-compatible server. docker run -d --ipc=host --device=/dev/neuron0 -p 8000:8080 docker.io/xbill9/gemma4-optb:slim inf2.xlarge optb slim : launch 2026-10-05T02:02:59+00:00 healthy 2026-10-05T02:17:18+00:00 - 14.3 min The graph is traced at a fixed size, 512 tokens in total and 128 for the prompt, and the server ignores tool definitions. A Strands agent's system prompt and tool schemas do not fit in 128 tokens, so the Neuron rows take part as tools of another agent Step 10 . Trainium1 and Inferentia2 report the same accelerator layout: two NeuronCore-v2 cores and 32 GiB of device memory per chip. The image compiled for inf2 runs on a trn1.2xlarge without a rebuild: READY in 121.2s — serving on :8080 SLIM host, peak RSS 19.53 GB req pt=17 ct=94 38.1tok/s 2.47s finish=stop req pt=18 ct=107 36.2tok/s 2.96s finish=stop The 26B mixture-of-experts build made for a single inf2.xlarge runs unchanged too: READY in 220.1s — slim int8-squeeze, ModelBuilder TP=2, MAX=512 BUCKET=128 req pt=21 ct=101 prefill=0.31s decode=5.5tok/s e2e=5.4tok/s 18.6s finish=length The difference is the host. A trn1.2xlarge has 8 vCPUs and 32 GiB of RAM against the inf2.xlarge 's 4 and 16, so the 19.53 GB peak while loading E2B fits in memory without a swapfile. The instances accept no inbound traffic. A port forward per instance puts each server on localhost : aws ssm start-session --region us-east-2 --target