On-Device Inference Debugging (Part 2): Threads on Big Cores, CPU at Full Clock — Still Slow A developer benchmarking Qwen2.5-1.5B on a Pixel 4 (Snapdragon 855) ruled out thread scheduling and CPU frequency throttling as the dominant cause of the phone's 0.5 tok/s inference speed, finding that forcing the performance governor via schedutil only improved generation by 20% (to 0.6 tok/s). Thread dumps confirmed all four inference threads stayed on the A76 mid and big cores (cpu4–7) with OpenMP working as configured, and battery temperature of 33.9 °C ruled out thermal throttling. The only confirmed fix so far remains adding -march for ARM cross-compilation, which cut prefill from 114 s to 45 s. Part 1: Why Your Phone Runs LLMs 80x Slower Than It Should This part focuses on thread scheduling and CPU frequency scaling . In part 1 I benchmarked Qwen2.5-1.5B on a Pixel 4 Snapdragon 855 and measured 0.5 tok/s — 60-120x off the theoretical limits for that SoC. That investigation found one real bug: llama.cpp cross-compilation passes no -march by default , so quantized matmul falls back to a slow path. Fixing it cut prefill from 114 s to 45 s. But throughput was still 0.5 tok/s. The biggest hole was filled; the problem remained. After publishing on DEV, a reader offered two traps that look like a slow model but aren't: Scheduling: a llama.cpp server whose child processes were parked on efficiency cores under load — a 3-second transcription became 37 seconds. Same binary, same flags; the only difference was who launched it. Storage: running off a USB stick, the whole model is re-read at every start — 2.6 GB on a 34 MB/s drive means over a minute before the first token. And the line that stung: The number that looks like "the model is slow" is almost never the model. Both point the same way: maybe the problem isn't what is being computed, but where and how fast . for c in /sys/devices/system/cpu/cpu 0-9 ; do echo "$ basename $c $ cat $c/cpufreq/cpuinfo max freq " done | Cores | Max freq | Type | |---|---|---| | cpu0–3 | 1.79 GHz | A55 little | | cpu4–6 | 2.42 GHz | A76 mid | | cpu7 | 2.84 GHz | A76 big | Classic 1+3+4. If inference threads land on cpu0–3, you lose 3-4x immediately. Field 39 of /proc/