cd /news/large-language-models/on-device-inference-debugging-part-2… · home › topics › large-language-models › article
[ARTICLE · art-148714] src=dev.to ↗ pub= topic=large-language-models verified=true sentiment=· neutral

On-Device Inference Debugging (Part 2): Threads on Big Cores, CPU at Full Clock — Still Slow

A developer benchmarking Qwen2.5-1.5B on a Pixel 4 (Snapdragon 855) ruled out thread scheduling and CPU frequency throttling as the dominant cause of the phone's 0.5 tok/s inference speed, finding that forcing the performance governor via schedutil only improved generation by 20% (to 0.6 tok/s). Thread dumps confirmed all four inference threads stayed on the A76 mid and big cores (cpu4–7) with OpenMP working as configured, and battery temperature of 33.9 °C ruled out thermal throttling. The only confirmed fix so far remains adding -march for ARM cross-compilation, which cut prefill from 114 s to 45 s.

by read5 min views2 publishedOct 10, 2026

Part 1: Why Your Phone Runs LLMs 80x Slower Than It Should

This part focuses on thread scheduling and CPU frequency scaling.

In part 1 I benchmarked Qwen2.5-1.5B on a Pixel 4 (Snapdragon 855) and measured 0.5 tok/s — 60-120x off the theoretical limits for that SoC.

That investigation found one real bug: llama.cpp cross-compilation passes no -march by default, so quantized matmul falls back to a slow path. Fixing it cut prefill from 114 s to 45 s.

But throughput was still 0.5 tok/s. The biggest hole was filled; the problem remained.

After publishing on DEV, a reader offered two traps that look like a slow model but aren't:

Scheduling: a llama.cpp server whose child processes were parked on efficiency cores under load — a 3-second transcription became 37 seconds. Same binary, same flags; the only difference was who launched it.

Storage: running off a USB stick, the whole model is re-read at every start — 2.6 GB on a 34 MB/s drive means over a minute before the first token.

And the line that stung:

The number that looks like "the model is slow" is almost never the model.

Both point the same way: maybe the problem isn't what is being computed, but where and how fast.

for c in /sys/devices/system/cpu/cpu[0-9]*; do
  echo "$(basename $c) $(cat $c/cpufreq/cpuinfo_max_freq)"
done
Cores Max freq Type
cpu0–3 1.79 GHz A55 little
cpu4–6 2.42 GHz A76 mid
cpu7 2.84 GHz A76 big

Classic 1+3+4. If inference threads land on cpu0–3, you lose 3-4x immediately.

Field 39 of /proc/<pid>/task/<tid>/stat is the CPU the thread is currently running on.

pid=$(pidof com.example.localai)

for t in /proc/$pid/task/*; do
  tid=$(basename $t)
  comm=$(cat $t/comm)
  cpu=$(awk '{print $39}' $t/stat)
  ticks=$(awk '{print $14+$15}' $t/stat)
  echo "$tid $comm cpu=$cpu ticks=$ticks"
done

Sample repeatedly — a single snapshot tells you nothing.

main thread 9556:  cpu5 -> cpu5 -> cpu5 -> cpu3 -> cpu5 -> cpu4
19524 openmp_worker  cpu6
19525 openmp_worker  cpu7   <- big core
19526 openmp_worker  cpu5

All four inference threads stayed on cpu4–7. None touched the little cores.

Also confirmed OpenMP is genuinely working:

19524 openmp_worker   5607 ticks
19525 openmp_worker   5606 ticks
19526 openmp_worker   5602 ticks
9556  DefaultDispatch 12132 ticks (main)

n_threads=4 → 3 workers + 1 main thread. Exactly as configured.

Verdict: "threads parked on efficiency cores" does not happen on this device. Ruled out.

cpu4:  710 ~ 1920 MHz   (max 2419)   <- as low as 29%
cpu5: 1612 ~ 2419 MHz   (max 2419)
cpu6:  710 ~ 2419 MHz   (max 2419)
cpu7: 1612 ~ 2016 MHz   (max 2841)   <- never above 71%
php
adb shell dumpsys battery | grep temperature

33.9 °C. Not thermal.

The culprit is schedutil, Android's default governor. It scales frequency based on utilization estimates — and LLM inference is bursty: small work units, back to back. The estimate lags, the frequency hunts, and you get exactly this oscillation.

for c in 0 1 2 3 4 5 6 7; do
  g=/sys/devices/system/cpu/cpu$c/cpufreq/scaling_governor
  chmod 666 "$g"
  echo performance > "$g"
done

Confirmed pinned: cpu0-3 @1785, cpu4-6 @2419, cpu7 @2841 MHz.

Metric schedutil performance Delta
Generation 0.5 tok/s 0.6 tok/s +20%
First token 1863 ms 1619 ms −13%

20%. Throttling is real, but it is not the dominant factor.

Hypothesis Verdict
Missing ARM optimization ( -march ) ✅ Real, fixed — 2.5x on prefill
Threads scheduled onto little cores ❌ Ruled out
OpenMP not actually threading ❌ Ruled out
mmap pages evicted ❌ Ruled out ( LLAMA_LOAD_MODE_MLOCK changed nothing)
CPU frequency throttling ⚠️ Real, but only ~20%
CPU at full clock (2419/2841 MHz)
threads on big cores (cpu4-7)
dotprod instructions present (verified by disassembly)
        |
        v
Utilization of peak compute capacity: ~1%

The arithmetic:

per token: 1.5B params x 2 = 3 GFLOPs
measured: 1.67 s per token
effective: 1.8 GFLOPS
peak: ~182 GFLOPS (4x A76 at full clock)
utilization: 1%

And memory:

per token reads all weights = 1.06 GB
1.06 GB / 1.67 s = 635 MB/s
available bandwidth: 17-34 GB/s
utilization: 2-4%

Neither the CPU nor the memory bus is anywhere near saturated. It's just slow.

The remaining explanation points inside ggml's quantized kernels on this platform — memory access patterns, cache behaviour, or OpenMP barrier overhead on small work units.

I can't fix that from configuration. It stays on the list as the next thing to dig into.

The investigation produced no final answer — but the process is reusable.

awk '{print $39}' /proc/<pid>/task/<tid>/stat      # current CPU per thread
awk '{print $14+$15}' /proc/<pid>/task/<tid>/stat  # cumulative CPU time per thread
cat /sys/devices/system/cpu/cpu*/cpufreq/scaling_cur_freq   # live frequency
cat /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor   # governor

Works on any Android device. No extra tooling.

Ruling out A, B and C is progress. It shrinks the search space from "anything" to "inside ggml's kernels".

If I only published the successful fix, you'd think this was a clean path. Real engineering is mostly wrong hypotheses and the value is eliminating them quickly.

Most people say "my model is slow" without quantifying it. Work out the memory-bandwidth ceiling and the compute ceiling, then measure. Knowing the gap is 60-120x tells you the problem isn't "needs tuning" — it's "something isn't running at all".

Item Restored to
CPU governor schedutil
SELinux Enforcing
stay-awake flag false
screen timeout 60000 ms
temp scripts deleted

Don't leave side effects on a test device — especially SELinux.

Next up: ggml quantized kernel efficiency on ARM.

If you've dug into this, I'd love to hear it.

Part 1: Why Your Phone Runs LLMs 80x Slower Than It Should

Part 2: this post

Code: https://github.com/Pingredsai/local-ai-android

── more in #large-language-models 4 stories · sorted by recency
── more on @qwen2.5-1.5b 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/on-device-inference-…] indexed:0 read:5min 2026-10-10 · —