22:14
2026-08-09
dev.to
artificial-intelligence
Self-hosting a lite agent backend on one TPU: Gemma 4 E2B + vLLM on a v5e-1
A developer self-hosted a lite agent backend on a single Google Cloud TPU v5e chip, achieving 1,496 output tokens/sec aggregate throughput with 8.02 ms per-token latency for Gemma 4 E2B under vLLM. Th…