17:59
2026-10-02
dev.to
large-language-models
Repacked QAT Gemma 4 on One TPU v5e: 12B Serves at 675 Tokens per Second
A developer repacked Google's quantization-aware-trained Gemma 4 weights for vLLM and served them on a single Google Cloud TPU v5e chip, publishing the scripts and per-record results on GitHub. The inβ¦