00:00
2026-08-20
nobodywho.ai
machine-learning
Use fewer threads for CPU inference
A benchmark of Gemma4-E4B on an AMD Ryzen 7040 CPU shows that using 8 threads instead of 16 improves prompt processing from 94.66 to 105.94 tokens per second and token generation from 11.64 to 16.15 t…