Benchmarking small LLM inference on SageMaker AI: G7 vs G5 and G6
Amazon Web Services benchmarked two 30B Mixture-of-Experts LLMs on SageMaker AI, showing that new G7 instances powered by NVIDIA Blackwell GPUs deliver measurable gains in throughput, latency, and cosβ¦