Running Gemma 4 26B at 5 tokens/SEC on a 13-year-old Xeon with no GPU
A developer running Google's Gemma 4 26B mixture-of-experts model on a 13-year-old HP StoreVirtual server with dual Xeon E5-2690 v2 CPUs and no GPU achieved approximately 5 tokens per second inference…