{"slug": "on-device-inference-debugging-part-2-threads-on-big-cores-cpu-at-full-clock-slow", "title": "On-Device Inference Debugging (Part 2): Threads on Big Cores, CPU at Full Clock — Still Slow", "summary": "A developer benchmarking Qwen2.5-1.5B on a Pixel 4 (Snapdragon 855) ruled out thread scheduling and CPU frequency throttling as the dominant cause of the phone's 0.5 tok/s inference speed, finding that forcing the performance governor via schedutil only improved generation by 20% (to 0.6 tok/s). Thread dumps confirmed all four inference threads stayed on the A76 mid and big cores (cpu4–7) with OpenMP working as configured, and battery temperature of 33.9 °C ruled out thermal throttling. The only confirmed fix so far remains adding -march for ARM cross-compilation, which cut prefill from 114 s to 45 s.", "body_md": "Part 1: *Why Your Phone Runs LLMs 80x Slower Than It Should*\n\nThis part focuses on **thread scheduling** and **CPU frequency scaling**.\n\nIn part 1 I benchmarked Qwen2.5-1.5B on a Pixel 4 (Snapdragon 855) and measured **0.5 tok/s** — 60-120x off the theoretical limits for that SoC.\n\nThat investigation found one real bug: **llama.cpp cross-compilation passes no `-march` by default**, so quantized matmul falls back to a slow path. Fixing it cut prefill from 114 s to 45 s.\n\n**But throughput was still 0.5 tok/s.** The biggest hole was filled; the problem remained.\n\nAfter publishing on DEV, a reader offered two traps that *look* like a slow model but aren't:\n\n**Scheduling:** a llama.cpp server whose child processes were parked on efficiency cores under load — a 3-second transcription became 37 seconds. Same binary, same flags; the only difference was who launched it.\n\n**Storage:** running off a USB stick, the whole model is re-read at every start — 2.6 GB on a 34 MB/s drive means over a minute before the first token.\n\nAnd the line that stung:\n\n**The number that looks like \"the model is slow\" is almost never the model.**\n\nBoth point the same way: maybe the problem isn't *what* is being computed, but *where* and *how fast*.\n\n```\nfor c in /sys/devices/system/cpu/cpu[0-9]*; do\n  echo \"$(basename $c) $(cat $c/cpufreq/cpuinfo_max_freq)\"\ndone\n```\n\n| Cores | Max freq | Type | \n|---|---|---|\n| cpu0–3 | 1.79 GHz | A55 little | \n| cpu4–6 | 2.42 GHz | A76 mid | \n| cpu7 | **2.84 GHz** | A76 big | \n\nClassic 1+3+4. If inference threads land on cpu0–3, you lose 3-4x immediately.\n\nField 39 of `/proc/<pid>/task/<tid>/stat` is **the CPU the thread is currently running on**.\n\n```\npid=$(pidof com.example.localai)\n\nfor t in /proc/$pid/task/*; do\n  tid=$(basename $t)\n  comm=$(cat $t/comm)\n  cpu=$(awk '{print $39}' $t/stat)\n  ticks=$(awk '{print $14+$15}' $t/stat)\n  echo \"$tid $comm cpu=$cpu ticks=$ticks\"\ndone\n```\n\n**Sample repeatedly** — a single snapshot tells you nothing.\n\n``` php\nmain thread 9556:  cpu5 -> cpu5 -> cpu5 -> cpu3 -> cpu5 -> cpu4\n19524 openmp_worker  cpu6\n19525 openmp_worker  cpu7   <- big core\n19526 openmp_worker  cpu5\n```\n\n**All four inference threads stayed on cpu4–7. None touched the little cores.**\n\nAlso confirmed OpenMP is genuinely working:\n\n```\n19524 openmp_worker   5607 ticks\n19525 openmp_worker   5606 ticks\n19526 openmp_worker   5602 ticks\n9556  DefaultDispatch 12132 ticks (main)\n```\n\n`n_threads=4` → 3 workers + 1 main thread. Exactly as configured.\n\n**Verdict: \"threads parked on efficiency cores\" does not happen on this device. Ruled out.**\n\n```\ncpu4:  710 ~ 1920 MHz   (max 2419)   <- as low as 29%\ncpu5: 1612 ~ 2419 MHz   (max 2419)\ncpu6:  710 ~ 2419 MHz   (max 2419)\ncpu7: 1612 ~ 2016 MHz   (max 2841)   <- never above 71%\nphp\nadb shell dumpsys battery | grep temperature\n# temperature: 339  ->  33.9 C\n```\n\n33.9 °C. Not thermal.\n\nThe culprit is **`schedutil`**, Android's default governor. It scales frequency based on utilization estimates — and LLM inference is *bursty*: small work units, back to back. The estimate lags, the frequency hunts, and you get exactly this oscillation.\n\n```\n# root required; SELinux Enforcing blocks a direct write, so chmod first\nfor c in 0 1 2 3 4 5 6 7; do\n  g=/sys/devices/system/cpu/cpu$c/cpufreq/scaling_governor\n  chmod 666 \"$g\"\n  echo performance > \"$g\"\ndone\n```\n\nConfirmed pinned: cpu0-3 @1785, cpu4-6 @2419, cpu7 @2841 MHz.\n\n| Metric | schedutil | performance | Delta | \n|---|---|---|---|\n| Generation | 0.5 tok/s | **0.6 tok/s** | **+20%** | \n| First token | 1863 ms | 1619 ms | −13% | \n\n**20%.** Throttling is real, but it is not the dominant factor.\n\n| Hypothesis | Verdict | \n|---|---|\n| Missing ARM optimization ( `-march` ) | ✅ **Real, fixed** — 2.5x on prefill | \n| Threads scheduled onto little cores | ❌ Ruled out | \n| OpenMP not actually threading | ❌ Ruled out | \n| mmap pages evicted | ❌ Ruled out ( `LLAMA_LOAD_MODE_MLOCK` changed nothing) | \n| CPU frequency throttling | ⚠️ Real, but only ~20% | \n\n```\nCPU at full clock (2419/2841 MHz)\nthreads on big cores (cpu4-7)\ndotprod instructions present (verified by disassembly)\n        |\n        v\nUtilization of peak compute capacity: ~1%\n```\n\nThe arithmetic:\n\n```\nper token: 1.5B params x 2 = 3 GFLOPs\nmeasured: 1.67 s per token\neffective: 1.8 GFLOPS\npeak: ~182 GFLOPS (4x A76 at full clock)\nutilization: 1%\n```\n\nAnd memory:\n\n```\nper token reads all weights = 1.06 GB\n1.06 GB / 1.67 s = 635 MB/s\navailable bandwidth: 17-34 GB/s\nutilization: 2-4%\n```\n\n**Neither the CPU nor the memory bus is anywhere near saturated. It's just slow.**\n\nThe remaining explanation points inside **ggml's quantized kernels on this platform** — memory access patterns, cache behaviour, or OpenMP barrier overhead on small work units.\n\n**I can't fix that from configuration. It stays on the list as the next thing to dig into.**\n\nThe investigation produced no final answer — but the *process* is reusable.\n\n```\nawk '{print $39}' /proc/<pid>/task/<tid>/stat      # current CPU per thread\nawk '{print $14+$15}' /proc/<pid>/task/<tid>/stat  # cumulative CPU time per thread\ncat /sys/devices/system/cpu/cpu*/cpufreq/scaling_cur_freq   # live frequency\ncat /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor   # governor\n```\n\n**Works on any Android device. No extra tooling.**\n\nRuling out A, B and C *is* progress. It shrinks the search space from \"anything\" to \"inside ggml's kernels\".\n\nIf I only published the successful fix, you'd think this was a clean path. Real engineering is mostly wrong hypotheses and the value is eliminating them quickly.\n\nMost people say \"my model is slow\" without quantifying it. Work out the memory-bandwidth ceiling and the compute ceiling, then measure. **Knowing the gap is 60-120x tells you the problem isn't \"needs tuning\" — it's \"something isn't running at all\".**\n\n| Item | Restored to | \n|---|---|\n| CPU governor | `schedutil` | \n| SELinux | `Enforcing` | \n| stay-awake flag | `false` | \n| screen timeout | 60000 ms | \n| temp scripts | deleted | \n\n**Don't leave side effects on a test device — especially SELinux.**\n\nNext up: **ggml quantized kernel efficiency on ARM**.\n\nIf you've dug into this, I'd love to hear it.\n\n*Part 1: [Why Your Phone Runs LLMs 80x Slower Than It Should](https://dev.to/pingredsai/why-your-phone-runs-llms-80x-slower-than-it-should-and-what-i-found-2645)*\n\n*Part 2: this post*\n\n*Code: [https://github.com/Pingredsai/local-ai-android](https://github.com/Pingredsai/local-ai-android)*", "url": "https://wpnews.pro/news/on-device-inference-debugging-part-2-threads-on-big-cores-cpu-at-full-clock-slow", "canonical_source": "https://dev.to/pingredsai/on-device-inference-debugging-part-2-threads-on-big-cores-cpu-at-full-clock-still-slow-6hm", "published_at": "2026-10-10 11:02:50+00:00", "updated_at": "2026-10-10 11:13:27.914539+00:00", "lang": "en", "topics": ["large-language-models", "ai-infrastructure", "mlops"], "entities": ["Qwen2.5-1.5B", "Pixel 4", "Snapdragon 855", "llama.cpp", "OpenMP", "Android", "schedutil"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/on-device-inference-debugging-part-2-threads-on-big-cores-cpu-at-full-clock-slow", "markdown": "https://wpnews.pro/news/on-device-inference-debugging-part-2-threads-on-big-cores-cpu-at-full-clock-slow.md", "text": "https://wpnews.pro/news/on-device-inference-debugging-part-2-threads-on-big-cores-cpu-at-full-clock-slow.txt", "jsonld": "https://wpnews.pro/news/on-device-inference-debugging-part-2-threads-on-big-cores-cpu-at-full-clock-slow.jsonld"}}