cd /news/ai-infrastructure/decode-decoded · home › topics › ai-infrastructure › article
[ARTICLE · art-147043] src=zhebrak.io ↗ pub= topic=ai-infrastructure verified=true sentiment=· neutral

Decode Decoded

A benchmark of vLLM 0.30.0, SGLang 0.5.20, and TensorRT-LLM 1.2.1 on a single H100 SXM GPU with 8 CPU cores found no host-bound regime across Qwen 3 0.6B-32B models at batch sizes from 1 to 256, with CUDA graphs enabled, according to the decode-decoded analysis published with code at zhebrak/decode-decoded. TensorRT-LLM underperformed the other engines when serving smaller models, which the author attributes to its custom attention kernel, assessed via Nsight Systems kernel name-matching, kernel counts per step, and time per step. CUDA graphs delivered the largest engine-level speedup, disproportionately for smaller models and smaller batch sizes, while scheduler overlap also reduced step time and vLLM's Model Runner V2 made no meaningful difference in these runs.

by read2 min views1 publishedOct 7, 2026
Decode Decoded
Image: source

Decode performance of vLLM, SGLang, and TensorRT-LLM on H100 SXM across Qwen 3 0.6B-32B models for 1-256 batch sizes analysed with Nsight Systems.

Somewhat surprisingly, the benchmarks didn’t record a host-bound regime even for the smallest model, meaning the GPU was always busy and the CPU never became the bottleneck (at least with CUDA graphs turned on). However, being busy does not always mean being efficient or fast, as the GPU can spend time on kernels, copies, or CUDA graph launches without fully utilising its memory bandwidth. Each step below is split into the bandwidth floor (reading weights and the KV cache from HBM at the card’s measured memory bandwidth), kernel excess (kernel time over the bandwidth floor), and kernel launch cost, with host stalling and copies as the remainder.

Less surprisingly, with smaller models in general and smaller batch sizes for smaller models, kernels’ fixed cost couldn’t be amortised by the bandwidth floor. Even with CUDA graphs, the inter-kernel launch overhead still exists, now within the card itself.

With the same cuBLAS kernels, the biggest difference between engines was when serving smaller models, with TensorRT-LLM underperforming because of its custom attention kernel (assessed via kernel name-matching of Nsight traces, kernel counts per step, and time per step). Hollow markers below are within the noise.

For engine optimisations, CUDA graphs made the biggest difference overall and are disproportionately more important for smaller models and smaller batch sizes in general. Scheduler overlap also reduces step time, rescuing performance for smaller models, although its relationship with batch size varied by engine. Model Runner V2 for vLLM made no meaningful difference in these runs. This benchmark ran on a host with one H100 SXM GPU and 8 CPU cores, on dense bf16 models, and is tied to specific releases of the serving engines (vLLM 0.30.0, SGLang 0.5.20, TensorRT-LLM 1.2.1). Prompts of 896 tokens; 288 tokens generated. Three timed generations per cell after two warm-up generations. Noise measured separately with representative runs. Code is available at zhebrak/decode_decoded.

── more in #ai-infrastructure 4 stories · sorted by recency
── more on @vllm 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/decode-decoded] indexed:0 read:2min 2026-10-07 · —