Gemma 4 E2B on a Single TPU v6e Chip: A Serving Deep Dive
Google's Gemma 4 E2B model serves efficiently on a single TPU v6e chip, achieving 213 tok/s for a single user and scaling to ~2,200 output tok/s across concurrent streams, while its QAT variants fail …