18:06
2026-10-07
zhebrak.io
ai-infrastructure
Decode Decoded
A benchmark of vLLM 0.30.0, SGLang 0.5.20, and TensorRT-LLM 1.2.1 on a single H100 SXM GPU with 8 CPU cores found no host-bound regime across Qwen 3 0.6B-32B models at batch sizes from 1 to 256, with …