Speculative Decoding in Llama.cpp: How to Actually Speed Up Local LLMs
Testing on a Panther Lake mini PC showed that speculative decoding with a draft model more than doubled local LLM token generation in Llama.cpp, from about 5 tokens per second to roughly 11 on an integrated GPU, accordin…