21:38
2026-10-02
justine.lol
large-language-models
LLaMA Now Goes Faster on CPUs
Justine Tunney wrote 84 new matrix multiplication kernels for llamafile that make prompt evaluation 30% to 500% faster than llama.cpp on CPU with F16 and Q8_0 weights, according to benchmarks publisheβ¦