{"slug": "llama-now-goes-faster-on-cpus", "title": "LLaMA Now Goes Faster on CPUs", "summary": "Justine Tunney wrote 84 new matrix multiplication kernels for llamafile that make prompt evaluation 30% to 500% faster than llama.cpp on CPU with F16 and Q8_0 weights, according to benchmarks published March 31, 2024. On a Skylake HP Intel Core i9-9900, llamafile-0.7 reached 28 prompt tok/sec on Mistral 7b q4_0 versus 17 for llama.cpp 2024-03-26 and 12 for llamafile-0.6.2, and the new kernels run 2x faster than Intel MKL for matrices that fit in L2 cache. The gains are largest on ARMv8.2+ (Raspberry Pi 5), Intel Alder Lake, and AVX512 (Zen 4) CPUs, and are currently limited to prompts under 1,000 tokens.", "body_md": "Mar 31<sup>st</sup>, 2024 @ [justine's web page](../index.html)\n\nI just wrote 84 new matrix multiplication kernels for\n[llamafile](https://github.com/mozilla-Ocho/llamafile) which\nenable it to read prompts / images faster. Compared to llama.cpp, prompt\neval time with llamafile should go anywhere between 30% and 500% faster\nwhen using F16 and Q8_0 weights on CPU. The improvements are most\ndramatic for ARMv8.2+ (e.g. RPI 5), Intel (e.g. Alderlake), and AVX512\n(e.g. Zen 4) computers. My kernels go 2x faster than MKL for matrices\nthat fit in L2 cache, which makes them a work in progress, since the\nspeedup works best for prompts having fewer than 1,000 tokens.\n\nllamafile is a local LLM project I started with Mozilla back in Nov\n2023. We're using\n[Cosmopolitan Libc](https://justine.lol/cosmo3/) to package\n[llama.cpp](https://github.com/ggerganov/llama.cpp/) as a\nsingle-file cross-platform binary that runs on six OSes for AMD64 and\nARM64, while making gentle modifications. I believe that by improving\nthe core technology, we can give our users the best possible llama.cpp\nexperience, while helping both projects reach a broader audience.\nMozilla has been giving me the resources to do this.\n\nWhen I first got into LLMs, my workstation was an austere Hewlett\nPackard running Alpine with a spinning disk, slow RAM, an AVX2\nprocessor, and no GPU. What I liked about llama.cpp is they were the\nfirst LLM project that cared about people like me. So I started\nvolunteering full time and collaborated with guys like Slaren to\n[introduce\nmmap()](https://github.com/ggerganov/llama.cpp/pull/613) support, which made weights load instantly using half as much\nRAM. It was a leap forward for local LLMs at the time, but did little to\nimprove evaluation speed. Most of the inference code was written by\nGeorgi Gerganov himself, and it's so good that it'd take me another year\nto finally improve upon. Now that I have, let's see how much faster\nthings go on my old Hewlett Packard.\n\n| LLM Performance on HP Intel® Core™ i9-9900 ($439) w/ 2200 MT/s RAM c. 2020 |  |  |  |  |  | \n|---|---|---|---|---|---|\n| prompt tok/sec | eval tok/sec | model | weights data type | hardware | software | \n| 28 | 7 | Mistral 7b | q4_0 | Skylake | llamafile-0.7 | \n| 17 | 7 | Mistral 7b | q4_0 | Skylake | llama.cpp 2024-03-26 | \n| 12 | 7 | Mistral 7b | q4_0 | Skylake | llamafile-0.6.2 | \n| 32 | 4 | Mistral 7b | q8_0 | Skylake | llamafile-0.7 | \n| 22 | 4 | Mistral 7b | q8_0 | Skylake | llama.cpp 2024-03-26 | \n| 16 | 4 | Mistral 7b | q8_0 | Skylake | llamafile-0.6.2 | \n| 23 | 2 | Mistral 7b | f16 | Skylake | llamafile-0.7 | \n| 15 | 2 | Mistral 7b | f16 | Skylake | llama.cpp 2024-03-26 | \n| 14 | 2 | Mistral 7b | f16 | Skylake | llamafile-0.6.2 | \n| 205 | 26 | TinyLlama 1.1B | q8_0 | Skylake | llamafile-0.7 | \n| 144 | 26 | TinyLlama 1.1B | q8_0 | Skylake | llama.cpp 2024-03-26 | \n| 91 | 23 | TinyLlama 1.1B | q8_0 | Skylake | llamafile-0.6.2 | \n| 171 | 15 | TinyLlama 1.1B | f16 | Skylake | llamafile-0.7 | \n| 118 | 15 | TinyLlama 1.1B | f16 | Skylake | llama.cpp 2024-03-26 | \n| 101 | 15 | TinyLlama 1.1B | f16 | Skylake | llamafile-0.6.2 | \n\nHere we see that, on Skylake, llamafile users can expect to see a 2x speedup and llama.cpp users can expect 50% better performance. Please note this only applies to certain weights. So far, I've only written optimized kernels for the q8_0, f16, q4_1, q4_0, and f32 data types. I think both q8_0 and f16 are really solid choices. Possibly even f32 if you've got plenty of RAM. That's because my new kernels change the rules. They're doing such a good job fixing the memory bandwidth quants always solved, that quantization could become the bigger bottleck. That would be great news for the future of local language models, since it means less need to trade away knowledge for speed.\n\nYou don't need a large computer to run a large language model. One of the best personal computers available in stores today is the Raspberry Pi. They deliver good performance at a great price and consume very little power.\n\n|  LLM Performance on $136 [Raspberry Pi v5](https://amzn.to/3TYRti8) (ARMv8.2) and v4 (ARMv8.0) |  |  |  |  |  | \n|---|---|---|---|---|---|\n| prompt tok/sec | eval tok/sec | model | weights data type | hardware | software | \n| 62 | 5 | TinyLlama 1.1B | f16 | RPI5 | llamafile-0.7 | \n| 28 | 5 | TinyLlama 1.1B | f16 | RPI5 | llama.cpp 2024-03-26 | \n| 8 | 5 | TinyLlama 1.1B | f16 | RPI5 | llamafile-0.6.2 | \n| 45 | 9 | TinyLlama 1.1B | q8_0 | RPI5 | llamafile-0.7 | \n| 35 | 9 | TinyLlama 1.1B | q8_0 | RPI5 | llama.cpp 2024-03-26 | \n| 20 | 9 | TinyLlama 1.1B | q8_0 | RPI5 | llamafile-0.6.2 | \n| 10 | 3 | TinyLlama 1.1B | q8_0 | RPI4 | llamafile-0.7 | \n| 10 | 3 | TinyLlama 1.1B | q8_0 | RPI4 | llama.cpp 2024-03-26 | \n| 9 | 3 | TinyLlama 1.1B | q8_0 | RPI4 | llamafile-0.6.2 | \n| 3 | 2 | TinyLlama 1.1B | f16 | RPI4 | llamafile-0.7 | \n| 3 | 2 | TinyLlama 1.1B | f16 | RPI4 | llama.cpp 2024-03-26 | \n| 4 | 2 | TinyLlama 1.1B | f16 | RPI4 | llamafile-0.6.2 | \n\nRaspberry Pi released their fifth edition a few months ago and it's outrageously fast compared to their previous model. They also introduced support for the ARMv8.2 dotprod and fp16 arithmetic ISAs, which are very useful for LLMs. Those two features alone enabled llama.cpp to achieve a 10x performance boost for f16 weights last year. This week I teased out another 2x performance boost on top of that, by using a kernel that I originally intended for AVX512. You wouldn't think a kernel designed for beefy data center equipment would work out for a teensy tiny little Raspberry Pi, but it actually fit hand in glove since both CPUs have 32 vector registers.\n\nIt's worthwhile to note that the new ARMv8.2 fp16 ISA may introduce more\nerrors than usual, since it causes llamafile to use fp16 words and we\naren't using techniques like [Kahan summation](kahan.gif) for\ncomputing dot products. So Q8_0 weights actually end up having slightly\nbetter perplexity, because it uses the dotprod ISA which lets us updot\nsigned 8-bit integers into a 32-bit compute type which absorbs errors.\nHowever this doesn't mean the faster fp16 weights can't be useful. Many\ndevelopers in this field view the differences as negligible.\n\nFor example, let's say you want to setup an email server on your pihole\nand have TinyLLaMA filter spam. It's possible to\n[configure Postfix\nto filter content using a shell script](https://www.postfix.org/FILTER_README.html) which lets you run the\nllamafile command.\n\n```\nllamafile -m TinyLlama-1.1B-Chat-v1.0.f16.gguf \\\n          --grammar 'root ::= \"yes\" | \"no\"' --temp 0 -c 0 \\\n          --no-display-prompt --log-disable -p \"<|user|>\nCan you say for certain that the following email is spam?\n\nTo: jtunney@gmail.com\nFrom: Federal-Tax-DebtHelp <ConfirmationEmail.inzk@janents.com>\nSubject: Reduce your payments to what you can afford\n\nReduce your payments to what you can afford \n \n [IMG] \n [IMG] \n \n [IMG] \n</s>\n<|assistant|>\"\n```\n\nWhen I run the shell script above on my RPI5, it takes 3 seconds.\n\n``` bash\njart@pi5:~/scratch$ time ./spam.sh\nyes\n\nreal    0m3.168s\nuser    0m10.851s\nsys     0m0.541s\n```\n\nThere are several important things happening here:\n\n`--temp 0` turns off the random number generator (we\n    don't want improvisation for a spam filter)\n  `--grammar` flag to force the LLM to only print a\n    single \"yes\\n\" or \"no\\n\" token.\n  ```\nlinks -codepage utf-8\n    -force-html -width 400 -dump /dev/stdin\n```\n which reduced the\n    number of tokens and removed hidden HTML content that was put there\n    by the spammer to make the email look like ham to naive Bayesian\n    filters.\n  `-c 0` flag configures TinyLLaMA to use the maximum\n    context size, which is 2048 tokens. That's the largest prompt we can\n    give it. To help avoid feeding in too much text, you can pipe the\n    email through ```\nsed 's/   */ /g' | dd bs=1\n    count=7000\n```\n to remove superfluous spaces and place an upper\n    limit on its size.\nPlease note that spam filtering is just the tip of the iceberg. I've always thought that \"generative ai\" is a misnomer, because language models (more commonly known as \"The Algorithm\") have always been used by tech companies in the past to extract knowledge, curate information, and rank content. That requires reading rather than writing. AI has never traditionally needed to talk that much, because the English language just isn't that useful at the scale tech companies operate. Even if you could generate English summaries for exabytes of text a millionth its size, that would still be more words than any team could hope to read in a lifetime.\n\nGamers have the highest quality expectations of any value consumer, so any hardware built for them is usually pretty good. In the machine learning industry, we have thrived for years repurposing hardware that was intended for gamers. If it weren't for their important contribution, the AI Winter may have needed to last another ten years. So a few months ago, I asked a gamer to build me a computer that can replace my old Hewlett Packard.\n\n|  LLM Performance on [Intel® Core™     i9-14900K](https://amzn.to/3I75E2e) ($438) w/ 6400 MT/s RAM |  |  |  |  |  | \n|---|---|---|---|---|---|\n| prompt tok/sec | eval tok/sec | model | weights data type | hardware | software | \n| 63 | 12 | Mistral 7b | q8_0 | Alderlake | llamafile-0.7 | \n| 40 | 9 | Mistral 7b | q8_0 | Alderlake | llama.cpp 2024-03-26 | \n| 19 | 7 | Mistral 7b | q8_0 | Alderlake | llamafile-0.6.2 | \n| 50 | 7 | Mistral 7b | f16 | Alderlake | llamafile-0.7 | \n| 13 | 5 | Mistral 7b | f16 | Alderlake | llama.cpp 2024-03-26 | \n| 10 | 4 | Mistral 7b | f16 | Alderlake | llamafile-0.6.2 | \n| 406 | 67 | TinyLlama 1.1B | q8_0 | Alderlake | llamafile-0.7 | \n| 273 | 53 | TinyLlama 1.1B | q8_0 | Alderlake | llama.cpp 2024-03-26 | \n| 114 | 43 | TinyLlama 1.1B | q8_0 | Alderlake | llamafile-0.6.2 | \n| 407 | 42 | TinyLlama 1.1B | f16 | Alderlake | llamafile-0.7 | \n| 90 | 31 | TinyLlama 1.1B | f16 | Alderlake | llama.cpp 2024-03-26 | \n| 68 | 30 | TinyLlama 1.1B | f16 | Alderlake | llamafile-0.6.2 | \n\nI think Alderlake is a great CPU but it's popularly misunderstood, as\nevidenced by how easily I quintupled its float16 performance. Unlike\nARMv8.2, I was able to do that without introducing rounding errors,\nsince my x86 kernels use a float32 compute type internally. This means I\ncan have an even smarter spam filter. For example, when I run\nmy `spam.sh` shell script, it only takes 420 milliseconds,\nwhich is 7x faster than my Raspberry Pi 5. That's right, when it comes\nto small workloads, this chip is able to finish before CUDA even gets\nstarted.\n\nAlderlake owners can also look forward to the fact that llamafile takes\nspecial care to not run on your efficiency cores. This is one of the\nthings that helps llamafile to go faster than llama.cpp. It also means\nyou can run LLMs around the clock and there's still plenty of resources\nleftover for the other programs on your computer. The reason why that's\nimportant is because llama.cpp dispatches threads in lockstep, which\nwould have meant that if any `1` core takes longer than the\nothers to do its job, then all other `n` cores would need to\nbusy loop until it completed.\n\nHowever the greatest feature of this microprocessor is how quickly it\ncan build all 2.6 million lines of code in the Cosmopolitan monorepo. My\nHewlett Packard always took 64 seconds, but this gaming computer does it\nin 20. It actually took 35 seconds originally; what made it faster was applying\n[liquid\nmetal](https://www.thermal-grizzly.com/en/conductonaut/s-tg-c-001-r) and AI overclocking. Another reason systems code is so fast on\nthe Alderlake is there was a fire fight between the hackers and\nscientists in the creation of this CPU, and the hackers won. I hope\nthey'll strike out a better compromise on AVX512 in the future, but\noverall I'm very happy with this chip, since I believe it represents\nsignificant progress over previous models.\n\nIf there's a personal computer with the most class, it would definitely be the Mac Studio. Gaining the performance advantage here was harder for me, because it's the hardware platform the llama.cpp developers care about most, plus I'm working with a handicap due to my choice to use Stallman's compiler instead of Apple's proprietary tools.\n\n| LLM Performance on Mac Studio CPU w/ 24-core M2 Ultra ($5000) |  |  |  |  |  | \n|---|---|---|---|---|---|\n| prompt tok/sec | eval tok/sec | model | weights data type | hardware | software | \n| 90 | 25 | Mistral 7b | q8_0 | M2 Ultra | llamafile-0.7 | \n| 90 | 27 | Mistral 7b | q8_0 | M2 Ultra | llama.cpp 2024-03-26 | \n| 37 | 24 | Mistral 7b | q8_0 | M2 Ultra | llamafile-0.6.2 | \n| 79 | 15 | Mistral 7b | f16 | M2 Ultra | llamafile-0.7 | \n| 57 | 15 | Mistral 7b | f16 | M2 Ultra | llama.cpp 2024-03-26 | \n| 21 | 15 | Mistral 7b | f16 | M2 Ultra | llamafile-0.6.2 | \n| 457 | 95 | TinyLlama 1.1B | q8_0 | M2 Ultra | llamafile-0.7 | \n| 564 | 108 | TinyLlama 1.1B | q8_0 | M2 Ultra | llama.cpp 2024-03-26 | \n| 236 | 95 | TinyLlama 1.1B | q8_0 | M2 Ultra | llamafile-0.6.2 | \n| 419 | 66 | TinyLlama 1.1B | f16 | M2 Ultra | llamafile-0.7 | \n| 400 | 67 | TinyLlama 1.1B | f16 | M2 Ultra | llama.cpp 2024-03-26 | \n| 141 | 66 | TinyLlama 1.1B | f16 | M2 Ultra | llamafile-0.6.2 | \n\nI wouldn't want to pick a fight with an Apple user, because their M2\nmicroprocessor turns llamafile into a firehose of synthetic content. The\ntrick Apple used to do it is leveraging their vertical integration. If\nyou buy a Mac Studio and look inside, you'll discover that they put the\nRAM DIMMs *inside* the CPU. It makes latency-bound operations\nlike token generation go much faster, because the CPU no longer needs to\nmake all these long distance phone calls. However, in terms of sheer\nflops (as measured by prompt tok/sec), we can see that compared to my\nmuch cheaper Intel computer, the M2 Ultra only exposes 30% more compute\nvia the ARM ISA. You need to go through their proprietary frameworks\nlike Metal and Accelerate if you want to access anything more. If you\nhave xcode installed, then llamafile by default will compile a small\nstub module which does just that, since despite my values I'm happy to\nhelp you get in front of any closed source library standing between you\nand your silicon.\n\nOne important thing to know if you're considering buying a Mac Studio is\nthat, like the Windows Executive, XNU does a really good job keeping\nyour desktop stable, and that means protecting your system from you. It\ntakes me 45 seconds on Mac Studio to compile the Cosmo monorepo, due to\nall these safety features; but if I fork bombed it, I'd be surprised if\nNetflix skipped a single frame. My `spam.sh` script also goes\n430ms, which is slower than Intel. However none of this concerns me,\nsince I've seen the way Asahi Linux is able to unleash the M2's full\npotential.\n\nWhile llamafile cares deeply about helping the GPU poor, it offers a first-class experience to the 1% too. The AMD Ryzen Threadripper PRO 7995WX was just launched several months ago and it's the most expensive CPU money can buy right now. It'll set you back $10,000 but you get 96 cores of AVX512, based on the Zen4 architecture.\n\n|      LLM Performance on [AMD Ryzen Threadripper PRO 7995WX](../rseq/threadripper.html) w/ 96 cores ($10,000) |  |  |  |  |  | \n|---|---|---|---|---|---|\n| prompt tok/sec | eval tok/sec | model | weights data type | hardware | software | \n| 557 | 17 | Mistral 7b | bf16 | 7995WX | llamafile-0.7 | \n| 485 | 17 | Mistral 7b | f16 | 7995WX | llamafile-0.7 | \n| 197 | 16 | Mistral 7b | f16 | 7995WX | llama.cpp 2024-03-29 | \n| 52 | 18 | Mistral 7b | f16 | 7995WX | llamafile-0.6.2 | \n| 480 | 10 | Mistral 7b | f32 | 7995WX | llamafile-0.7 | \n| 221 | 10 | Mistral 7b | f32 | 7995WX | llama.cpp 2024-03-30 | \n| 38 | 9 | Mistral 7b | f32 | 7995WX | llamafile-0.6.2 | \n| 382 | 25 | Mistral 7b | q8_0 | 7995WX | llamafile-0.7 | \n| 283 | 24 | Mistral 7b | q8_0 | 7995WX | llama.cpp 2024-03-29 | \n| 37 | 25 | Mistral 7b | q8_0 | 7995WX | llamafile-0.6.2 | \n| 1929 | 52 | TinyLlama 1.1B | bf16 | 7995WX | llamafile-0.7 | \n| 1819 | 52 | TinyLlama 1.1B | f16 | 7995WX | llamafile-0.7 | \n| 824 | 51 | TinyLlama 1.1B | f16 | 7995WX | llama.cpp 2024-03-29 | \n| 295 | 89 | TinyLlama 1.1B | f16 | 7995WX | llamafile-0.6.2 | \n| 1268 | 60 | TinyLlama 1.1B | q8_0 | 7995WX | llamafile-0.7 | \n| 1127 | 60 | TinyLlama 1.1B | q8_0 | 7995WX | llama.cpp 2024-03-29 | \n| 169 | 93 | TinyLlama 1.1B | q8_0 | 7995WX | llamafile-0.6.2 | \n\nHere we see that, despite only being twice the price, the 7995WX x86 ISA offers 7x more raw compute power than the M2 Ultra ARM ISA, and nearly the same token generation speed, which is likely thanks to its 384mb L3 cache. When I bought this chip, I had to expand support in llama.cpp for bfloat16 and AVX512 before I could fully test its capabilities. My work means you can now run LLaMA 2.8x faster on Zen4 than you could before.\n\nOne thing I like about AVX512 is that Google's Gemma model can\n[solve math\nriddles on AVX512 but not on AVX2](https://github.com/google/gemma.cpp/issues/23) because the bigger vectors usually\nmake it easier to reduce rounding errors. Its `VDPBF16PS`\ninstruction helps us updot bf16 similar to VNNI and ARM dotprod. Having\nnative support for bf16 is nice, since models like Mistral and TinyLLaMA\ndistribute weights using bfloat16 as their canonical format. If we were\nto convert bf16 to fp16, then only 13% of the numbers that are possible\ncan be accurately represented. In practice, it matters little, since\n99.71% of the numbers Mistral 7b uses are among that 13%. However I\nbelieve that llamafile should deliver, to the best of its ability,\nwhatever number of bits are being claimed. Especially when doing so also\nenables us to better exploit the capabilities of our hardware. Adding\nbf16 support is my first big step towards improving that.\n\nPlease be warned that a lot of people who bought this Threadripper ran\ninto issues with sketchy RAM. I had to RMA the first DIMMs I bought for\nthis computer, because most of them died and I was getting 5 eval tokens\nper second with Mistral. I've been having better luck with\na [new full kit of eight sticks](https://amzn.to/40EiMSH)\nthat just arrived today. When I run `sysbench memory run` it\nreports 10,033,424 mops, which is oddly faster than my Mac Studio where\n9,892,584 mops is reported, however my Intel computer does 14,490,952. I\nexpected my Threadripper's RAM to have that speed since both set of\ncomponents advertised 6400 MT/s with the same timings, but I'm told that\nI traded this away to have 256GB of ECC. As for disk speed, ```\ndd\nif=/dev/zero of=/tmp/output bs=128k count=50k; rm -f /tmp/output\n```\nreports 1.6 GB/s which is 3.6x slower than my Mac Studio, and 3x slower\nthan my Intel (which has the same M.2 stick). I'm told that Intel and\nApple are just better at this, but I wish I understood why.\n\nLast but not least, it runs my `spam.sh` script in 323ms and\nbuilds the whole Cosmo monorepo in 13 seconds. It's actually capable of\nbuilding it faster, since this is the first time I've ever seen my build\nconfig being constrained by an individual artifact blocking the critical\npath. I never thought I'd live to see the day. I'm also puzzled that\nllamafile v0.6.2 is somehow managing to do 93 tokens per second; that's\n40% faster than my M2. It's exciting news, since after reviewing the\nbreadth of this blog post, I would have wept if there were no more\noptimizations possible.\n\nThe source code for my matrix multiplication kernels can be found at:\n\nBoth Mozilla and myself felt it would be worthwhile to contribute these improvements to the upstream project. Doing that required adapting the code to their preferred method of handling portability at compile-time. We also took the liberty of changing the license from Apache 2.0 to MIT, since the latter is what the llama.cpp developers prefer. Here are links to the most recent pull requests I've sent them:\n\nThere are dozens of mathematical operations a transformer model needs to\nperform in order to generate text, e.g. rope, transpose, reshape,\nsoftmax, rms_norm, etc. All of the performance improvements I described\nabove, were achieved by focusing exclusively on a single one, which\nis `GGML_OP_MUL_MAT`, because that's what my Linux Perf\nprofiler told me llamafile spends 95% of its time doing.\n\nSo what is this matrix multiplication thing? We shall start by defining the most important algorithm in the world using the pythonic dialect of Python that is most popular with developers today:\n\n``` python\ndef matmul(A, B):\n  assert len(B) == len(A[0])\n  return [[sum(A[i][l] * B[l][j]\n               for l in range(len(B)))\n           for j in range(len(B[0]))]\n          for i in range(len(A))]\n```\n\nAs we can see, it's just three for loops and a multiply-add. How hard could it be?\n\nOn my workstation (which I call meatball), the code above goes a\nscreeching 0.042 gigaflops. Most Python programmers are smart enough to\nknow that they should delegate tasks like these to a library\nlike `np.matmul`, which goes 29 gigaflops. NumPy achieves its\nspeed using FORTRAN which for generations has been favored\nby [real programmers](https://justine.lol/dox/pascal.txt)\nwho've led us to believe these libraries are something mysterious\nchiseled in stone by the hand of Moses himself; but if we look at the\nFORTRAN code NumPy actually uses, then it really isn't all that\ncomplicated and could clearly benefit from some revision.\n\n```\n      SUBROUTINE SGEMM(TRANSA,TRANSB,M,N,K,ALPHA,A,LDA,B,LDB,BETA,C,LDC)\n*     .. Scalar Arguments ..\n      REAL ALPHA,BETA\n      INTEGER K,LDA,LDB,LDC,M,N\n      CHARACTER TRANSA,TRANSB\n*     .. Array Arguments ..\n      REAL A(LDA,*),B(LDB,*),C(LDC,*)\n      [...]\n*\n*           Form  C := alpha*A*B + beta*C.\n*\n              DO 90 J = 1,N\n                  IF (BETA.EQ.ZERO) THEN\n                      DO 50 I = 1,M\n                          C(I,J) = ZERO\n   50                 CONTINUE\n                  ELSE IF (BETA.NE.ONE) THEN\n                      DO 60 I = 1,M\n                          C(I,J) = BETA*C(I,J)\n   60                 CONTINUE\n                  END IF\n                  DO 80 L = 1,K\n                      IF (B(L,J).NE.ZERO) THEN\n                          TEMP = ALPHA*B(L,J)\n                          DO 70 I = 1,M\n                              C(I,J) = C(I,J) + TEMP*A(I,L)\n   70                     CONTINUE\n                      END IF\n   80             CONTINUE\n   90         CONTINUE\n      [...]\n      RETURN\n      END\n```\n\nI like to define my subroutines using a modern language like C++, which goes 47 gigaflops. This means C++ is three orders of a magnitude faster than Python. That's twenty years of progress per Moore's law.\n\n```\n// multiplies matrices on cpu\n// with column major ordering\n//\n//     m×k * k×n → m×n\n//     k×m * k×n → m×n if aᵀ\n//     m×k * n×k → m×n if bᵀ\n//     k×m * n×k → m×n if aᵀ and bᵀ\n//\ntemplate <typename T,  typename TA,\n          typename TB, typename TC>\nvoid GEMM(bool aᵀ, bool bᵀ,\n          int m, int n, int k, T α,\n          const TA *A, int lda,\n          const TB *B, int ldb, T β,\n          TC *C, int ldc) {\n    assert(m >= 0 && n >= 0 && k >= 0);\n    assert(lda >= std::max(1, aᵀ ? k : m));\n    assert(ldb >= std::max(1, bᵀ ? n : k));\n    assert(ldc >= std::max(1, m));\n#pragma omp parallel for collapse(2) if (m * n * k > 300000)\n    for (int i = 0; i < m; ++i)\n        for (int j = 0; j < n; ++j) {\n            T d = 0;\n            for (int l = 0; l < k; ++l) {\n                T a = A[aᵀ ? lda * i + l : lda * l + i];\n                T b = B[bᵀ ? ldb * l + j : ldb * j + l];\n                d += a * b;\n            }\n            if (β) {\n                T c = C[ldc * j + i];\n                C[ldc * j + i] = α * d + β * c;\n            } else {\n                C[ldc * j + i] = α * d;\n            }\n        }\n}\n```\n\nIn order to do better than 47 gigaflops on CPU, most C++ developers are\nsmart enough to know they should use a BLAS library. Mightiest of the\nopen source BLAS is [BLIS](https://github.com/flame/blis/)\nwhich is funded by Microsoft, Intel, Texas Instruments, AMD, HPE,\nOracle, Huawei, Facebook, ARM, and the National Science Foundation.\n\n\"Any time somebody outside Intel beats MKL by a nontrivial amount, I\nreport it to the MKL team. It is fantastic for any open-source project\nto get within 10% of MKL... [T]his is why Intel funds BLIS development.\"\n(@jeffhammond) [blis/issues/264](https://github.com/flame/blis/issues/264#issuecomment-428673275)\n\nThat's very impressive. Matrix multiplication is the practical\napplication of hardware that hardware makers care about optimizing most.\nSince nobody knows more about Intel hardware than Intel, I imagine it's\nnot everday that somebody manages to challenge Intel for supremacy on\ntheir own platform. Based on my own evaluation, what BLIS says is true.\nHowever that is only true for single-threaded performance. Their\nmultithreading mode is still experimental, but if I use\na `./configure` flag to turn it on, then I'm able to boost\nperformance to 85 gigaflops.\n\nllama.cpp had the important insight that less is more when it comes to linear algebra. The alpha and beta parameters are never used, so they're always set to to 1 and 0. The op graph for LLMs are designed in such a way that the A matrix is almost always transposed and B is almost never transposed, which means inner dimension dot product can vectorize over contiguous memory. The m/k dimensions are usually evenly divisible by 64. While generating tokens, n=1 is usually the case, which makes matmul a de facto matvec for the performance most people care about. BLAS libraries usually hurt more than they help for matrix-vector multiplication, because it's so computationally simple by comparison. Sort of like the difference between downloading a movie and pinging a server. Matrix vector multiplication is an operation where latency (not throughput) is the bottleneck, and the bloat of fancy libraries has a measurable impact. So llama.cpp does something like this, which goes 233 gigaflops.\n\n``` js\ntemplate <typename T>\nvoid LLMM(int m, int n, int k,\n          const T *A, int lda,\n          const T *B, int ldb,\n          T *C, int ldc) {\n#pragma omp parallel for collapse(2) if (m * n * k > 300000)\n    for (int i = 0; i < m; ++i)\n        for (int j = 0; j < n; ++j) {\n            T d = 0;\n            for (int l = 0; l < k; ++l)\n                d += A[lda * i + l] * B[ldb * j + l];\n            C[ldc * j + i] = d;\n        }\n}\n```\n\nThis gives us the best possible token generation speeds. However llama.cpp's Achilles heel on CPU has always been prompt processing speed, which goes much slower. That's because chewing through prompts requires bona fide matrix-matrix multiplication. Being able to do this fast is important if you care about text summarization and LLaVA image processing. That's the reason why support for countless BLAS libraries has been added to llama.cpp over the past year. The most formidable of them is Intel's Math Kernel Library (MKL) which goes 384 gigaflops.\n\nThe difference between 233 versus 384 gigaflops may not seem like much, at least not compared to Python, but it's a tremendous gulf. MKL is also closed source and proprietary. We aren't even allowed to disassemble it and try to reverse engineer how it works. Intel has been developing math kernels for fifty years and they hold the secrets they've acquired very close to their chest. But even if our desire for performance was so great that we were willing to overlook the ethics of an open source project spending the majority of its time inside a proprietary blob, the simple fact of the matter is that integrating foreign BLAS libraries into llama.cpp isn't that practical, due to the way its threading model works. In order to improve prompt processing speed, we must figure out the trick BLAS libraries use, and implement it in a scrappy dependency-free way that stays true to llama.cpp's roots.\n\nI believe the trick with CPU math kernels is exploiting instruction\nlevel parallelism with fewer memory references. If you compile the\nexample above with `-O3 -ffast-math -march=native` then the\ncode your compiler generates should look like this:\n\n``` js\nvoid SLLMM(int m, int n, int k,\n           const float *A, int lda,\n           const float *B, int ldb,\n           float *C, int ldc) {\n#pragma omp parallel for collapse(2) if (m * n * k > 300000)\n    for (int i = 0; i < m; ++i)\n        for (int j = 0; j < n; ++j) {\n            __m256 c = _mm256_setzero_ps();\n            for (int l = 0; l < k; l += 8)\n                c = _mm256_fmadd_ps(_mm256_loadu_ps(A + lda * i + l),\n                                    _mm256_loadu_ps(B + ldb * j + l), c);\n            C[ldc * j + i] = hsum(c);\n        }\n}\n```\n\nSo what llama.cpp usually does when it wants to improve things, is it'll unroll the innermost loop like this:\n\n``` js\nvoid SLLMM2(int m, int n, int k,\n           const float *A, int lda,\n           const float *B, int ldb,\n           float *C, int ldc) {\n#pragma omp parallel for collapse(2) if (m * n * k > 300000)\n    for (int i = 0; i < m; ++i)\n        for (int j = 0; j < n; ++j) {\n            __m256 c0 = _mm256_setzero_ps();\n            __m256 c1 = _mm256_setzero_ps();\n            for (int l = 0; l < k; l += 16) {\n                c0 = _mm256_fmadd_ps(_mm256_loadu_ps(A + lda * i + l + 0),\n                                     _mm256_loadu_ps(B + ldb * j + l + 0), c0);\n                c1 = _mm256_fmadd_ps(_mm256_loadu_ps(A + lda * i + l + 8),\n                                     _mm256_loadu_ps(B + ldb * j + l + 8), c1);\n            }\n            C[ldc * j + i] = hsum(c0) + hsum(c1);\n        }\n}\n```\n\nThat may slightly improve numerical stability, but it does very little\nto enhance performance, since modern CPUs are perfectly capable of\nspeculatively executing future loop iterations on their own. What we\nwant to do instead is unroll the *outer* loop. The advantage of\ndoing this becomes clear if we consider how it enables us to share\nthe `a0` register load across multiple floating point\noperations.\n\n``` js\nvoid SLLMM4(int m, int n, int k,\n            const float *A, int lda,\n            const float *B, int ldb,\n            float *C, int ldc) {\n#pragma omp parallel for collapse(2) if (m * n * k > 300000)\n    for (int i = 0; i < m; ++i)\n        for (int j = 0; j < n; j += 4) {\n            __m256 c0 = _mm256_setzero_ps();\n            __m256 c1 = _mm256_setzero_ps();\n            __m256 c2 = _mm256_setzero_ps();\n            __m256 c3 = _mm256_setzero_ps();\n            for (int l = 0; l < k; l += 8) {\n                __m256 a0 = _mm256_loadu_ps(A + lda * (i + 0) + l);\n                __m256 k0 = _mm256_loadu_ps(B + ldb * (j + 0) + l);\n                __m256 k1 = _mm256_loadu_ps(B + ldb * (j + 1) + l);\n                __m256 k2 = _mm256_loadu_ps(B + ldb * (j + 2) + l);\n                __m256 k3 = _mm256_loadu_ps(B + ldb * (j + 3) + l);\n                c0 = _mm256_fmadd_ps(a0, k0, c0);\n                c1 = _mm256_fmadd_ps(a0, k1, c1);\n                c2 = _mm256_fmadd_ps(a0, k2, c2);\n                c3 = _mm256_fmadd_ps(a0, k3, c3);\n            }\n            C[ldc * (j + 0) + (i + 0)] = hsum(c0);\n            C[ldc * (j + 1) + (i + 0)] = hsum(c1);\n            C[ldc * (j + 2) + (i + 0)] = hsum(c2);\n            C[ldc * (j + 3) + (i + 0)] = hsum(c3);\n        }\n}\n```\n\nIf we unroll both outer loops, the effect is compounded.\n\n``` js\nvoid SLLMM3X4(int m, int n, int k,\n              const float *A, int lda,\n              const float *B, int ldb,\n              float *C, int ldc) {\n#pragma omp parallel for collapse(2) if (m * n * k > 300000)\n    for (int i = 0; i < m; i += 3)\n        for (int j = 0; j < n; j += 4) {\n            __m256 c00 = _mm256_setzero_ps();\n            __m256 c01 = _mm256_setzero_ps();\n            __m256 c02 = _mm256_setzero_ps();\n            __m256 c03 = _mm256_setzero_ps();\n            __m256 c10 = _mm256_setzero_ps();\n            __m256 c11 = _mm256_setzero_ps();\n            __m256 c12 = _mm256_setzero_ps();\n            __m256 c13 = _mm256_setzero_ps();\n            __m256 c20 = _mm256_setzero_ps();\n            __m256 c21 = _mm256_setzero_ps();\n            __m256 c22 = _mm256_setzero_ps();\n            __m256 c23 = _mm256_setzero_ps();\n            for (int l = 0; l < k; l += 8) {\n                __m256 k0 = _mm256_loadu_ps(B + ldb * (j + 0) + l);\n                __m256 k1 = _mm256_loadu_ps(B + ldb * (j + 1) + l);\n                __m256 k2 = _mm256_loadu_ps(B + ldb * (j + 2) + l);\n                __m256 k3 = _mm256_loadu_ps(B + ldb * (j + 3) + l);\n                __m256 a0 = _mm256_loadu_ps(A + lda * (i + 0) + l);\n                c00 = _mm256_fmadd_ps(a0, k0, c00);\n                c01 = _mm256_fmadd_ps(a0, k1, c01);\n                c02 = _mm256_fmadd_ps(a0, k2, c02);\n                c03 = _mm256_fmadd_ps(a0, k3, c03);\n                __m256 a1 = _mm256_loadu_ps(A + lda * (i + 1) + l);\n                c10 = _mm256_fmadd_ps(a1, k0, c10);\n                c11 = _mm256_fmadd_ps(a1, k1, c11);\n                c12 = _mm256_fmadd_ps(a1, k2, c12);\n                c13 = _mm256_fmadd_ps(a1, k3, c13);\n                __m256 a2 = _mm256_loadu_ps(A + lda * (i + 2) + l);\n                c20 = _mm256_fmadd_ps(a2, k0, c20);\n                c21 = _mm256_fmadd_ps(a2, k1, c21);\n                c22 = _mm256_fmadd_ps(a2, k2, c22);\n                c23 = _mm256_fmadd_ps(a2, k3, c23);\n            }\n            C[ldc * (j + 0) + (i + 0)] = hsum(c00);\n            C[ldc * (j + 1) + (i + 0)] = hsum(c01);\n            C[ldc * (j + 2) + (i + 0)] = hsum(c02);\n            C[ldc * (j + 3) + (i + 0)] = hsum(c03);\n            C[ldc * (j + 0) + (i + 1)] = hsum(c10);\n            C[ldc * (j + 1) + (i + 1)] = hsum(c11);\n            C[ldc * (j + 2) + (i + 1)] = hsum(c12);\n            C[ldc * (j + 3) + (i + 1)] = hsum(c13);\n            C[ldc * (j + 0) + (i + 2)] = hsum(c20);\n            C[ldc * (j + 1) + (i + 2)] = hsum(c21);\n            C[ldc * (j + 2) + (i + 2)] = hsum(c22);\n            C[ldc * (j + 3) + (i + 2)] = hsum(c23);\n        }\n}\n```\n\nVectorized outer product with OpenMP goes 810 gigaflops on my Alderlake\ni9-14900K with 6400 MT/s RAM when multiplying a 513×512 with a 512×512\nmatrix. That is twenty eight years of progress per Moore's law compared\nto Python. It's clearly optimal since my CPU is listed as only being\ncapable of going\n[780\ngigaflops](https://nanoreview.net/en/cpu/intel-core-i9-14900k). Yes, I overclocked it with liquid metal. On the other\nhand, MKL processes this matrix size at 295 gigaflops on my machine.\n\n```\n1:        vmovups      (%r10,%r9,4),%ymm0\n          vmovups      (%rsi,%r9,4),%ymm4\n          vmovups      (%rcx,%r9,4),%ymm2\n          vmovups      (%rdx,%r9,4),%ymm1\n          vfmadd231ps  (%r11,%r9,4),%ymm0,%ymm6\n          vfmadd231ps  %ymm4,%ymm0,%ymm15\n          vfmadd231ps  %ymm2,%ymm0,%ymm12\n          vfmadd231ps  %ymm1,%ymm0,%ymm9\n          vmovups      (%rdi,%r9,4),%ymm0\n          vfmadd231ps  (%r11,%r9,4),%ymm0,%ymm5\n          vfmadd231ps  %ymm4,%ymm0,%ymm14\n          vfmadd231ps  %ymm2,%ymm0,%ymm11\n          vfmadd231ps  %ymm1,%ymm0,%ymm8\n          vmovups      (%rbx,%r9,4),%ymm0\n          vfmadd231ps  (%r11,%r9,4),%ymm0,%ymm3\n          add          $8,%r9\n          vfmadd231ps  %ymm4,%ymm0,%ymm13\n          vfmadd231ps  %ymm2,%ymm0,%ymm10\n          vfmadd231ps  %ymm1,%ymm0,%ymm7\n          cmp          %r9d,%r14d\n          jg          1b\n```\n\nBut does the C function above generalize to all matrix sizes? Nope. If I bump the complexity up from 512 to 1024, then I'm pretty much back at square one, not doing much better than a naive kernel, and MKL wins once more. I personally don't view this as too problematic, since llama.cpp by default processes prompts in modestly sized batches, and a kernel should only need to be good for its intended size. It's also only a matter of time until I unriddle the tricks needed for optimal tiling and cache locality that can make my kernels scale.\n\nNow to incorporate this into llamafile, we can't use OpenMP for the same\nreason we can't use BLAS libraries. The kernel must be harmonized with\nthe way llama.cpp works. Its threading model is very similar to GPUs.\nOps in the model graph are processed one by one. A thread is spawned for\neach core. Threads are restrained by a spinlock barrier and then set\nloose to compute different parts of an output matrix in parallel as soon\nas the next op is ready for execution. The id of each thread is\ncalled `ith` and the number of threads is\ncalled `nth`. There are no futexes or semaphores, because\nkernel scheduling would greatly reduce tokens/sec. If we were to have\nthe `ith=0` thread call a BLAS API that spawned threads of\nits own, then they'd be immediately starved of resources by all\nthe `ith>0` threads returning to the spinlock barrier. We\ncan work within this model by defining a new kernel framework.\n\n```\n#define BEGIN_KERNEL(RM, RN) \\\n    int ytiles = (m - m0) / RM; \\\n    int xtiles = (n - n0) / RN; \\\n    int tiles = ytiles * xtiles; \\\n    int duty = (tiles + nth - 1) / nth; \\\n    int start = duty * ith; \\\n    int end = start + duty; \\\n    if (end > tiles) \\\n        end = tiles; \\\n    for (int job = start; job < end; ++job) { \\\n        int i = m0 + job / xtiles * RM; \\\n        int j = n0 + job % xtiles * RN;\n\n#define END_KERNEL() }\n```\n\nAlong with a solution for packing tiles.\n\n``` js\ntemplate <typename T> class GEMMER {\n  public:\n    GEMMER(int k, const T *A, int lda, const T *B, int ldb, float *C, int ldc,\n           int ith, int nth)\n        : k(k), A(A), lda(lda), B(B), ldb(ldb), C(C), ldc(ldc), ith(ith), nth(nth) {\n    }\n\n    void llmm(int m, int n) {\n        mnpack(0, m, 0, n);\n    }\n\n  private:\n    void mnpack(int m0, int m, int n0, int n) {\n        if (m - m0 <= 0 || n - n0 <= 0)\n            return;\n        int mc, nc, mp, np;\n        if (m - m0 >= 3 && n - n0 >= 4) {\n            mc = 3;\n            nc = 4;\n            llmm3x4(m0, m, n0, n);\n        } else if (m - m0 >= 4 && n - n0 >= 1) {\n            mc = 4;\n            nc = 1;\n            llmm4x1(m0, m, n0, n);\n        } else if (m - m0 >= 1 && n - n0 >= 4) {\n            mc = 1;\n            nc = 4;\n            llmm1x4(m0, m, n0, n);\n        } else {\n            mc = 1;\n            nc = 1;\n            llmm1x1(m0, m, n0, n);\n        }\n        mp = m0 + (m - m0) / mc * mc;\n        np = n0 + (n - n0) / nc * nc;\n        mnpack(mp, m, n0, np);\n        mnpack(m0, mp, np, n);\n        mnpack(mp, m, np, n);\n    }\n\n    // ...\n\n    void llmm1x1(int m0, int m, int n0, int n) {\n        BEGIN_KERNEL(1, 1)\n        __m256 c = _mm256_setzero_ps();\n        for (int l = 0; l < k; l += 8)\n            c = _mm256_fmadd_ps(_mm256_loadu_ps(A + lda * i + l),\n                                _mm256_loadu_ps(B + ldb * j + l), c);\n        C[ldc * j + i] = hsum(c);\n        END_KERNEL()\n    }\n\n    const int k;\n    const T *const A;\n    const int lda;\n    const T *const B;\n    const int ldb;\n    float *const C;\n    const int ldc;\n    const int ith;\n    const int nth;\n};\n```\n\nWe can now export nice friendly C APIs to GGML that go 790 gigaflops while incurring none of the latency disadvantages associated with traditional BLAS libraries.\n\n``` js\nvoid SLLMMT(int m, int n, int k,\n            const float *A, int lda,\n            const float *B, int ldb,\n            float *C, int ldc,\n            int ith, int nth) {\n    if (nth) {\n        GEMMER<float> tb{k, A, lda, B, ldb, C, ldc, ith, nth};\n        tb.llmm(m, n);\n    } else if (!HAVE_OPENMP || n * m * k < THRESHOLD) {\n        GEMMER<float> tb{k, A, lda, B, ldb, C, ldc, 0, 1};\n        tb.llmm(m, n);\n    } else {\n        nth = sysconf(_SC_NPROCESSORS_ONLN);\n#pragma omp parallel for\n        for (ith = 0; ith < nth; ++ith) {\n            GEMMER<float> tb{k, A, lda, B, ldb, C, ldc, ith, nth};\n            tb.llmm(m, n);\n        }\n    }\n}\n```\n\nYou need to run the following command on Linux in order to benchmark llamafile reliably. It also helps a little bit with timings to run as root, but that shouldn't be necessary.\n\n```\necho performance | sudo tee /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor\n```\n\nOn Apple Silicon, I needed to build llama.cpp using the following command in order to get it to run in CPU mode.\n\n```\nmake -j32 LLAMA_NO_ACCELERATE=1 LLAMA_NO_METAL=1\n```\n\nSome of you might be surprised that I didn't write my kernels in\nassembly like BLIS does, especially given my history of projects\nlike [blink](github.com/jart/blink), [SectorLISP](../sectorlisp2/)\nand [SectorLAMBDA](../sectorlambda2/). The truth is I've been\ncoding in assembly this whole time. I configured Emacs so I can push a\nbutton, and the disassembly for the C++ code I'm working on will pop up\non the screen in a few milliseconds. I know anyone whose codebase has\nslow build times doesn't possess this advantage, which has made me\nfamous. Once I figure out how to do that for .cu files, I'll be\nunstoppable.\n\nI learned how to write math kernels by renting\n[Vast](https://vast.ai/) VMs and watching\n[Gautham Venkatasubramanian](https://ahgamut.github.io/)\nand\n[mrdomino](https://github.com/mrdomino) develop CUDA kernels\nin a tmux session. They've been focusing on solving a much more\nimportant challenge for llamafile, which is helping it not have a\nmandatory dependency on the cuBLAS: the reigning supreme linear algebra\nlibrary of such speed, accuracy, and ferocity that it could only have\nbeen written by the prince of darkness himself. You're encouraged to\nfollow our ongoing progress on GitHub. The monospace font used on this\npage is called\n[PragmataPro](https://fsd.it/shop/fonts/pragmatapro/) and it\nwas was designed\nby [Fabrizio\nSchiavi](https://en.wikipedia.org/wiki/Fabrizio_Schiavi) in Italy.\n\nMy full-time work on open source projects like llamafile is funded\nthanks to the generous support of Mozilla,\nmy [GitHub sponsors](https://github.com/sponsors/jart), and\n[Patreon subscribers](https://www.patreon.com/jart). Thank\nyou everyone, for helping me have the opportunity to serve you these\nlast four years. Your support made it possible for high-quality math\nkernels to be shared with the commons.", "url": "https://wpnews.pro/news/llama-now-goes-faster-on-cpus", "canonical_source": "http://justine.lol/matmul/", "published_at": "2026-10-02 21:38:02+00:00", "updated_at": "2026-10-02 22:06:22.160332+00:00", "lang": "en", "topics": ["large-language-models", "ai-infrastructure", "machine-learning", "developer-tools"], "entities": ["llamafile", "Mozilla", "llama.cpp", "Justine Tunney", "Cosmopolitan Libc", "Mistral 7b", "TinyLlama 1.1B", "Raspberry Pi"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/llama-now-goes-faster-on-cpus", "markdown": "https://wpnews.pro/news/llama-now-goes-faster-on-cpus.md", "text": "https://wpnews.pro/news/llama-now-goes-faster-on-cpus.txt", "jsonld": "https://wpnews.pro/news/llama-now-goes-faster-on-cpus.jsonld"}}