Gemma 4 on a Tesla T4: QAT Weights Decode 1.79x Faster Than bf16
A developer deployed Google's Gemma 4 E2B to a single Tesla T4 GPU on a Compute Engine VM using vLLM 0.29.0, finding that the QAT int4-weight checkpoint decodes at 72.31 tok/s per stream versus 40.44 …