g5g vs g6 for LLM Serving: the Same Code, and 3.7x the Throughput
A developer benchmarked AWS g5g and g6 GPU instances for serving Google's Gemma-4-E2B language model with identical code and weights, finding the g6 delivers 3.7x the decode throughput. Profiling reve…