LLaMA Now Goes Faster on CPUs Justine Tunney wrote 84 new matrix multiplication kernels for llamafile that make prompt evaluation 30% to 500% faster than llama.cpp on CPU with F16 and Q8_0 weights, according to benchmarks published March 31, 2024. On a Skylake HP Intel Core i9-9900, llamafile-0.7 reached 28 prompt tok/sec on Mistral 7b q4_0 versus 17 for llama.cpp 2024-03-26 and 12 for llamafile-0.6.2, and the new kernels run 2x faster than Intel MKL for matrices that fit in L2 cache. The gains are largest on ARMv8.2+ (Raspberry Pi 5), Intel Alder Lake, and AVX512 (Zen 4) CPUs, and are currently limited to prompts under 1,000 tokens. Mar 31