Gemma 4 on an Old 4 GB Laptop GPU: QAT Takes It From 9.5 GiB to 1.6
A developer deployed Google's Gemma 4 E2B quantization-aware-trained (QAT) checkpoint on a 4 GB GTX 1650 Ti laptop GPU, shrinking the model from 9.5 GiB in bfloat16 to a 3.35 GB Q4_0 GGUF file with only 1.31 GiB resident…