00:00
2026-08-20
runagentrun.co.uk
large-language-models
Dual 3090s: the bottleneck isn't the GPU
Two benchmarks of Qwen3.8-27B on a single RTX 3090 show a 3.2x performance gap: 41.49 tok/s with llama.cpp (build b10088) versus 132 tok/s with vLLM using a DFlash2 block drafter, according to Insiderβ¦