{"slug": "gemma-4-inference-on-aws-bedrock-sagemaker-gpus-inferentia-and-trainium-behind", "title": "Gemma 4 Inference on AWS: Bedrock, SageMaker, GPUs, Inferentia and Trainium Behind One Strands Agent", "summary": "A developer benchmarked six AWS LLM inference backends — Amazon Bedrock, a SageMaker real-time endpoint, vLLM on two EC2 GPU families, and a hand-ported Gemma 4 build on Inferentia2 and Trainium — all driven by a single Strands agent using the same prompts and grading scripts. The survey documents setup, failure modes and pricing per path, including capacity errors on inf2.xlarge in us-east-2a and Trainium spot quota requests that opened support cases in four US regions without automatic approval. The g6 vLLM launch took 14.0 minutes to become healthy, and the author notes Bedrock's inference profile id (us.amazon.nova-micro-v1:0) is required for on-demand throughput.", "body_md": "This article provides a step by step survey of LLM inference on AWS: Amazon Bedrock, a SageMaker real-time endpoint, vLLM on two EC2 GPU families, and a hand-ported Gemma 4 on AWS Inferentia2 and Trainium. A Strands agent drives all six backends, and every number below comes from two complete runs on the same day with the same prompts.\n\n[https://github.com/xbill9/aws-inference-strands](https://github.com/xbill9/aws-inference-strands)\n\nAWS offers a model as an API call, as a managed endpoint, as a GPU you rent by the hour, and as two families of its own accelerator chips. Each path has its own setup, its own failure modes and its own price, and most comparisons cover one or two of them.\n\nThis survey puts all of them behind one agent. The same Strands code talks to Bedrock, to SageMaker, to vLLM on EC2, and to a Neuron server, and the same scripts grade the answers and time the requests.\n\n| Backend | What serves the model | Model | \n|---|---|---|\n| Amazon Bedrock | managed API | `us.amazon.nova-micro-v1:0` | \n| EC2 g6.xlarge (NVIDIA L4) | vLLM 0.30 | `xbill9/gemma-4-E2B-it-qat-q4_0-w4a16-ct-text-emb4` | \n| SageMaker `ml.g6.xlarge` | AWS vLLM 0.30 container | the same repack | \n| EC2 g5g.2xlarge (Graviton2 + T4G) | vLLM 0.27.2rc0, patched for sm_75 | `google/gemma-4-E2B-it` , fp16 | \n| EC2 inf2.xlarge (Inferentia2) | hand-ported Gemma 4, `xbill9/gemma4-optb:slim` | Gemma 4 E2B, compiled for Neuron | \n| EC2 trn1.2xlarge (Trainium) | the same image, unchanged | the same build | \n\nThe g6 and SageMaker rows run the same weights on the same GPU, one as a VM and one as a managed endpoint. The inf2 and trn1 rows run the same compiled image on two different AWS chips.\n\n`AmazonSSMManagedInstanceCore`, and a security group with no inbound rules.` pip install 'strands-agents[openai,sagemaker]' boto3`.\nQuota is permission to ask. Capacity is whether AWS has the machine. Both decide whether a run happens today.\n\n```\nus-east-2   Running On-Demand Trn instances 8.0\nus-east-2   All Trn Spot Instance Requests  0.0\nus-east-2   Trn spot request: CASE_OPENED   8.0 2026-10-04T21:36:11.599000-04:00\n```\n\nEight vCPUs of on-demand Trainium buys exactly one `trn1.2xlarge`. Spot starts at zero. A request for eight vCPUs of spot in each of the four US regions opened a support case in each one, and none was approved automatically. The instance type itself is offered in one zone of the three regions checked: `us-east-2c`.\n\nAn `inf2.xlarge` launch in `us-east-2a`, with 80 vCPUs of Inferentia quota free, returned:\n\n```\nAn error occurred (InsufficientInstanceCapacity) when calling the RunInstances operation: We currently do not have sufficient inf2.xlarge capacity in the Availability Zone you requested (us-east-2a).\n```\n\nThe same launch in `us-east-2b` succeeded.\n\nBedrock is the reference point: no instance, no endpoint, an IAM permission and a model id. The inference profile id (`us.amazon.nova-micro-v1:0`) is the one to use; the bare model id returns a validation error for on-demand throughput.\n\n``` python\nfrom strands.models import BedrockModel\nmodel = BedrockModel(model_id=\"us.amazon.nova-micro-v1:0\", region_name=\"us-east-1\", temperature=0.0)\n```\n\nThe g6 serves the 4-bit embedding repack of Gemma 4 E2B with vLLM's OpenAI server. Tool calls need two flags; without them the agent's tools are ignored.\n\n```\ndocker run -d --gpus all --ipc=host -p 8000:8000 vllm/vllm-openai:v0.30.0 \\\n  --model xbill9/gemma-4-E2B-it-qat-q4_0-w4a16-ct-text-emb4 \\\n  --max-model-len 8192 --gpu-memory-utilization 0.90 \\\n  --enable-auto-tool-choice --tool-call-parser gemma4\ng6 (vLLM 0.30, emb4): launch 2026-10-05T02:02:48+00:00 healthy 2026-10-05T02:16:46+00:00 -> 14.0 min\n```\n\nStrands reaches it with the OpenAI provider:\n\n``` python\nfrom strands.models.openai import OpenAIModel\nmodel = OpenAIModel(client_args={\"base_url\": \"http://localhost:8001/v1\", \"api_key\": \"unused\"},\n                    model_id=\"xbill9/gemma-4-E2B-it-qat-q4_0-w4a16-ct-text-emb4\")\n```\n\nThe g5g pairs a Graviton2 Arm host with an NVIDIA T4G: Arm and CUDA in an instance that costs $0.556 an hour. vLLM's arm64 image ships without the T4G's sm_75 kernels, and Gemma 4's 512-wide attention heads need more shared memory than Turing has, so this row runs a vLLM 0.27.2rc0 built from source for sm_75 with a shared-memory patch, from a prebuilt AMI. The same two tool flags go on its command line.\n\n```\n(EngineCore pid=2128) INFO 10-05 02:55:22 [default_loader.py:430] Loading weights took 520.97 seconds\n(APIServer pid=1651) INFO 10-05 02:58:21 [parser_manager.py:37] \"auto\" tool choice has been enabled.\n```\n\nThe 521 seconds of weight loading come from the first read of a fresh volume restored from the AMI snapshot.\n\nThe SageMaker row uses the AWS vLLM container on `ml.g6.xlarge`, the same GPU as the EC2 g6, with vLLM flags passed as `SM_VLLM_*` environment variables:\n\n```\n{\n    \"SM_VLLM_ENABLE_AUTO_TOOL_CHOICE\": \"true\",\n    \"SM_VLLM_MODEL\": \"xbill9/gemma-4-E2B-it-qat-q4_0-w4a16-ct-text-emb4\",\n    \"SM_VLLM_TOOL_CALL_PARSER\": \"gemma4\"\n}\n# SageMaker emb4 endpoint: make deploy returned at epoch 1791166343, InService at epoch 1791166834 (aws sagemaker wait) -> 8.2 min\n```\n\nStrands reaches it with `SageMakerAIModel`, which signs `invoke_endpoint` calls instead of opening a URL:\n\n``` python\nfrom strands.models.sagemaker import SageMakerAIModel\nmodel = SageMakerAIModel(endpoint_config={\"endpoint_name\": \"gemma-4-e2b-emb4-strands\", \"region_name\": \"us-east-2\"},\n                         payload_config={\"max_tokens\": 512, \"stream\": False, \"temperature\": 0.0})\n```\n\nNo vLLM release serves Gemma 4 on Neuron, so this row runs a hand-ported Gemma 4: a `torch_neuronx` graph compiled ahead of time and baked into a container with an OpenAI-compatible server.\n\n```\ndocker run -d --ipc=host --device=/dev/neuron0 -p 8000:8080 docker.io/xbill9/gemma4-optb:slim\ninf2.xlarge (optb slim): launch 2026-10-05T02:02:59+00:00 healthy 2026-10-05T02:17:18+00:00 -> 14.3 min\n```\n\nThe graph is traced at a fixed size, 512 tokens in total and 128 for the prompt, and the server ignores tool definitions. A Strands agent's system prompt and tool schemas do not fit in 128 tokens, so the Neuron rows take part as tools of another agent (Step 10).\n\nTrainium1 and Inferentia2 report the same accelerator layout: two NeuronCore-v2 cores and 32 GiB of device memory per chip. The image compiled for inf2 runs on a `trn1.2xlarge` without a rebuild:\n\n```\nREADY in 121.2s — serving on :8080 (SLIM host, peak RSS 19.53 GB)\n[req] pt=17 ct=94 38.1tok/s 2.47s finish=stop\n[req] pt=18 ct=107 36.2tok/s 2.96s finish=stop\n```\n\nThe 26B mixture-of-experts build made for a single `inf2.xlarge` runs unchanged too:\n\n```\nREADY in 220.1s — slim int8-squeeze, ModelBuilder TP=2, MAX=512 BUCKET=128\n[req] pt=21 ct=101 prefill=0.31s decode=5.5tok/s e2e=5.4tok/s 18.6s finish=length\n```\n\nThe difference is the host. A `trn1.2xlarge` has 8 vCPUs and 32 GiB of RAM against the `inf2.xlarge`'s 4 and 16, so the 19.53 GB peak while loading E2B fits in memory without a swapfile.\n\nThe instances accept no inbound traffic. A port forward per instance puts each server on `localhost`:\n\n```\naws ssm start-session --region us-east-2 --target <instance-id> \\\n  --document-name AWS-StartPortForwardingSession \\\n  --parameters '{\"portNumber\":[\"8000\"],\"localPortNumber\":[\"8001\"]}'\n8001 200\n8002 200\n8003 200\n```\n\nThe same Strands agent asks how many of eleven ids are 10 or more. Its only tool, `count_ids(op, value)`, returns the exact count, minimum and maximum, so the model chooses a filter and quotes a number. A run is graded on the answer and on the filter it sent: an agent that sends `id > 10` gets 7 back, and saying 8 anyway is right by accident.\n\n```\npython3 demo1_counting.py --runs 20 --json results/demo1-engine.json\nbackend     runs  correct  right-filter  errors  median s  filters sent\nbedrock       20   20/20          20/20       0       1.1  id >= 10\ng6            20   20/20          20/20       0      0.41  id >= 10\nsagemaker     20   20/20          20/20       0      0.58  id >= 10\nbackend     runs  correct  right-filter  errors  median s  filters sent\ng5g           20   20/20          20/20       0      2.26  id >= 10\n```\n\nWith `--tools rows`, the tool returns the eleven ids and the model counts them itself. All four backends answered 8 in 20 of 20 runs.\n\nThe Neuron servers cannot drive an agent, but any server can answer a question. Here a Strands agent on Bedrock gets one tool per backend, `ask_<name>(prompt)`, plus `scoreboard()`. Each `ask_` tool times its own request in code; `scoreboard()` sorts and formats the table; the agent is told to quote it as is.\n\n```\npython3 demo2_orchestrator.py\n| Backend | Answer | Seconds | Tokens/s |\n|---------|--------|---------|----------|\n| vLLM on EC2 g6 (L4) | Paris | 0.154 | 13.0 |\n| Trainium (trn1, hand-ported) | Paris | 0.196 | 5.1 |\n| Inferentia2 (inf2, hand-ported) | Paris | 0.242 | 4.1 |\n| Amazon Bedrock | Paris | 0.51 | 3.9 |\n| SageMaker endpoint (vLLM, L4) | Paris | 0.569 | 3.5 |\n\nAll backends have correctly identified the capital of France as \"Paris\". The fastest response came from the Trainium (trn1, hand-ported) backend.\n\n[harness] 5 backend calls made; not asked by the agent: none\n```\n\nEvery backend answered in all three runs. In this run the table, computed in code, puts the g6 first, and the agent's own closing sentence names Trainium. The table is the result; the sentence is the model's reading of it.\n\nA one-word answer is too short to time decoding, so a separate script sends `\"Write a short paragraph about the ocean.\"` with a 100-token limit, five times per backend, timing each request in code:\n\n```\npython3 bench_chat.py --repeats 5 --json results/bench-chat.json\nbackend     median s  median tok/s  tok/s range     tokens\nbedrock        0.894         111.9    87.7-128.0    100-100\ng6             0.639         131.5   130.9-131.7    84-84\nsagemaker       0.79         106.3    92.3-110.5    84-84\ninf2           2.583          36.4    33.1-36.6     94-94\ntrn1           2.534          37.1    35.7-37.4     94-94\nbackend     median s  median tok/s  tok/s range     tokens\ng5g            2.607          37.2    36.9-37.3     97-97\n```\n\nTokens per second here includes the network path and the prompt, measured from the laptop through each port forward or AWS endpoint.\n\nThe second run brought every backend up from scripts in the repository's `stage/` directory, then ran every demo again at full size:\n\n```\nstage/up.sh 2>&1 | tee stage/up.log        # launch g6, inf2, trn1, g5g and the SageMaker endpoint\nstage/forward.sh && source stage/demo.env   # SSM port forwards and the backend variables\npython3 stage/check.py --wait 1500          # one answer and one tool call per backend\n```\n\n`up.sh` tries every zone in a region, then the next region. inf2 found no capacity anywhere in us-east-2 or us-east-1:\n\n```\nstrands-demo-inf2: no capacity in us-east-2 (subnet-0880f3d5ac599127d), trying the next zone\nstrands-demo-inf2: no capacity in us-east-1 (subnet-09e0f13c0a5b43092), trying the next zone\nstrands-demo-inf2: launched i-00eae995ffe871f80 us-west-2a\n```\n\nLaunch to the first passing check:\n\n| Backend | Launch → ready | \n|---|---|\n| SageMaker | 8.0 min to InService | \n| trn1 | 12.8 min | \n| g6 | 13.9 min | \n| inf2, us-west-2 | 17.9 min | \n| g5g | 23.7 min, of which `Loading weights took 530.83 seconds` | \n\nDemo 1 scored 20 of 20 with the filter `id >= 10` on all four tool-capable backends again, and 20 of 20 with the model counting the rows. Demo 2 ran five times; the agent asked all six backends every time, and all 30 answers were Paris. g6 was fastest in four runs and trn1 in one, by 9 ms, which is the kind of margin a model's closing sentence can misread.\n\nThe benchmark ran ten repeats per backend. Medians against the first run, computed in code:\n\n| Backend | First run tok/s | Second run tok/s | Change | \n|---|---|---|---|\n| g6 | 131.5 | 132.45 | +0.7% | \n| Bedrock | 111.9 | 115.85 | +3.5% | \n| SageMaker | 106.3 | 105.2 | −1.0% | \n| g5g | 37.2 | 35.8 | −3.8% | \n| trn1 | 37.1 | 36.9 | −0.5% | \n| inf2 | 36.4 | 36.6 | +0.5% | \n\nEvery self-hosted backend stayed within 4% of its first run, inf2 in a different region. At temperature 0 each self-hosted backend returned one completion across its ten repeats, and Bedrock returned 9 distinct completions in 10. **g6 and SageMaker returned the same text word for word**, the same weights on the same GPU type, and **so did inf2 and trn1**, one image digest on two chips in two regions.\n\n`stage/down.sh` does this in one command. Every instance, endpoint, endpoint configuration and model goes; the security groups and the instance profile stay for the next run. The check covers all four US regions, because a run that fails over to another zone or region leaves its resources there:\n\n```\n=== us-east-1 inst:[] vols:[] sm-ep:[] sm-cfg:[] sm-model:[] spot:[]\n=== us-east-2 inst:[] vols:[] sm-ep:[] sm-cfg:[] sm-model:[] spot:[]\n=== us-west-1 inst:[] vols:[] sm-ep:[] sm-cfg:[] sm-model:[] spot:[]\n=== us-west-2 inst:[] vols:[] sm-ep:[] sm-cfg:[] sm-model:[] spot:[]\n```\n\nQuota lets you ask; the zone decides. Check which zones offer the instance type (`aws ec2 describe-instance-type-offerings`) before launching, keep a second zone and a second region ready, and start anything slow, such as a SageMaker endpoint or a Neuron model load, well before you need it.\n\nOn the second run inf2 was refused in five zones across us-east-2 and us-east-1 and launched in us-west-2a; `stage/up.sh` does that walk for you.\n\nEvery path here was brought up with E2B first: it loads on Neuron in about two minutes, against close to four for the 26B. Wiring, flags, port forwards and tool parsing all fail the same way on a small model as on a large one, and a small model fails in minutes. The 26B went to Trainium after the E2B run had already answered.\n\nLet code count, time and rank, and have the model quote the result. Then check what the model sent as well as what it said: demo 1 grades the filter, and demo 2's agent once quoted a correct table and then summarised it wrong.\n\nOne instance profile with SSM access served every EC2 instance in this survey, and one SSM-only security group per region served every launch in it. Broad roles and reusable groups keep the next run to a single `run-instances` call; tighten them before anything faces users.\n\nStrands 1.55's SageMaker provider builds its usage record from four fields, and vLLM 0.30 also returns `completion_tokens_details`, which raises `TypeError` on every call. The repository's `backends.py` drops the extra fields before Strands reads them; the token counts are unchanged.\n\nvLLM-Neuron 0.24.0.1.1.0 and 0.21.0.1.0.0 list Trn2 and Trn3 only. The 0.5.3 line that still covers Trn1 and Inf2 pins vLLM 0.16. None of them lists Gemma, so Gemma 4 on Neuron means writing the model yourself, which is what the hand-ported image is.\n\nSingle-request decode and on-demand price, us-east-1 list prices. Cost per million tokens is arithmetic: the hourly price divided by the measured tokens per hour.\n\n| Backend | tok/s (median) | $/hr | $ per million tokens, one request | \n|---|---|---|---|\n| 🥇 EC2 g6.xlarge, vLLM | 131.5 | 0.8048 | 1.70 | \n| 🥈 EC2 g5g.2xlarge, vLLM | 37.2 | 0.556 | 4.15 | \n| 🥉 EC2 inf2.xlarge, hand-ported | 36.4 | 0.7582 | 5.79 | \n| EC2 trn1.2xlarge, hand-ported | 37.1 | 1.3438 | 10.06 | \n\n| Backend | Setup on the day | Tool calls | Price model | \n|---|---|---|---|\n| Bedrock | none | yes | per token | \n| SageMaker | 8.2 and 8.0 min to InService | yes, via `SM_VLLM_*` | per instance hour | \n| EC2 g6 | 14.0 and 13.9 min to healthy | yes | per instance hour | \n| EC2 g5g | patched vLLM on a prebuilt AMI, 23.7 min to ready | yes | per instance hour | \n| EC2 inf2 / trn1 | inf2 14.3 min, and 17.9 min in us-west-2; trn1 12.8 min | no | per instance hour | \n\nStart on Bedrock: nothing to run, tool calls work, and you pay per token. Move to vLLM on a g6 when you need your own weights, a specific repack, or a full GPU to yourself; it was the fastest and cheapest per token here, and the same container on SageMaker adds a managed endpoint for an hourly premium. The g5g is the Arm-plus-CUDA option, and it runs Gemma 4 only with a patched vLLM. Inferentia2 and Trainium run the same hand-ported image at the same speed, with Inferentia2 at a lower hourly price for this model size; choose them when the model already exists for Neuron, because vLLM does not provide it.\n\nThe goal of this article was to survey six ways to serve a model on AWS behind one agent. The key to the solution was one Strands agent with a model object per backend, grading and timing done in code, and a fixed prompt set on one day. The results were:\n\n`us-east-2a` had none on the first run, and on the second none in any zone of us-east-2 or us-east-1, so it ran in us-west-2.\nScope: two runs per backend on 2026-10-05, us-east-2 except the g5g in us-east-1a and the second run's inf2 in us-west-2a, on-demand instances, client on a laptop reaching EC2 through SSM port forwards. Twenty agent runs per backend in demo 1 in each run, three then five in demo 2, five then ten timed requests per backend in the benchmark. Four different models or builds were compared: Bedrock serves Nova Micro, the g6 and SageMaker rows serve the 4-bit embedding repack of Gemma 4 E2B, the g5g serves the stock E2B in fp16, and the Neuron rows serve the hand-ported E2B build, so the speed table compares serving paths with the model each path supports.\n\nThe strategy for serving Gemma 4 on AWS from Bedrock to Trainium behind one Strands agent was validated with an incremental step by step approach.", "url": "https://wpnews.pro/news/gemma-4-inference-on-aws-bedrock-sagemaker-gpus-inferentia-and-trainium-behind", "canonical_source": "https://dev.to/gde/gemma-4-inference-on-aws-bedrock-sagemaker-gpus-inferentia-and-trainium-behind-one-strands-agent-2lnm", "published_at": "2026-10-06 15:09:17+00:00", "updated_at": "2026-10-06 15:18:17.465120+00:00", "lang": "en", "topics": ["large-language-models", "ai-infrastructure", "ai-chips", "mlops", "ai-agents"], "entities": ["AWS", "Amazon Bedrock", "Amazon SageMaker", "AWS Inferentia2", "AWS Trainium", "vLLM", "Gemma 4", "Strands"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/gemma-4-inference-on-aws-bedrock-sagemaker-gpus-inferentia-and-trainium-behind", "markdown": "https://wpnews.pro/news/gemma-4-inference-on-aws-bedrock-sagemaker-gpus-inferentia-and-trainium-behind.md", "text": "https://wpnews.pro/news/gemma-4-inference-on-aws-bedrock-sagemaker-gpus-inferentia-and-trainium-behind.txt", "jsonld": "https://wpnews.pro/news/gemma-4-inference-on-aws-bedrock-sagemaker-gpus-inferentia-and-trainium-behind.jsonld"}}