Part 1: Why Your Phone Runs LLMs 80x Slower Than It Should
This part focuses on thread scheduling and CPU frequency scaling.
In part 1 I benchmarked Qwen2.5-1.5B on a Pixel 4 (Snapdragon 855) and measured 0.5 tok/s — 60-120x off the theoretical limits for that SoC.
That investigation found one real bug: llama.cpp cross-compilation passes no -march by default, so quantized matmul falls back to a slow path. Fixing it cut prefill from 114 s to 45 s.
But throughput was still 0.5 tok/s. The biggest hole was filled; the problem remained.
After publishing on DEV, a reader offered two traps that look like a slow model but aren't:
Scheduling: a llama.cpp server whose child processes were parked on efficiency cores under load — a 3-second transcription became 37 seconds. Same binary, same flags; the only difference was who launched it.
Storage: running off a USB stick, the whole model is re-read at every start — 2.6 GB on a 34 MB/s drive means over a minute before the first token.
And the line that stung:
The number that looks like "the model is slow" is almost never the model.
Both point the same way: maybe the problem isn't what is being computed, but where and how fast.
for c in /sys/devices/system/cpu/cpu[0-9]*; do
echo "$(basename $c) $(cat $c/cpufreq/cpuinfo_max_freq)"
done
| Cores | Max freq | Type |
|---|---|---|
| cpu0–3 | 1.79 GHz | A55 little |
| cpu4–6 | 2.42 GHz | A76 mid |
| cpu7 | 2.84 GHz | A76 big |
Classic 1+3+4. If inference threads land on cpu0–3, you lose 3-4x immediately.
Field 39 of /proc/<pid>/task/<tid>/stat is the CPU the thread is currently running on.
pid=$(pidof com.example.localai)
for t in /proc/$pid/task/*; do
tid=$(basename $t)
comm=$(cat $t/comm)
cpu=$(awk '{print $39}' $t/stat)
ticks=$(awk '{print $14+$15}' $t/stat)
echo "$tid $comm cpu=$cpu ticks=$ticks"
done
Sample repeatedly — a single snapshot tells you nothing.
main thread 9556: cpu5 -> cpu5 -> cpu5 -> cpu3 -> cpu5 -> cpu4
19524 openmp_worker cpu6
19525 openmp_worker cpu7 <- big core
19526 openmp_worker cpu5
All four inference threads stayed on cpu4–7. None touched the little cores.
Also confirmed OpenMP is genuinely working:
19524 openmp_worker 5607 ticks
19525 openmp_worker 5606 ticks
19526 openmp_worker 5602 ticks
9556 DefaultDispatch 12132 ticks (main)
n_threads=4 → 3 workers + 1 main thread. Exactly as configured.
Verdict: "threads parked on efficiency cores" does not happen on this device. Ruled out.
cpu4: 710 ~ 1920 MHz (max 2419) <- as low as 29%
cpu5: 1612 ~ 2419 MHz (max 2419)
cpu6: 710 ~ 2419 MHz (max 2419)
cpu7: 1612 ~ 2016 MHz (max 2841) <- never above 71%
php
adb shell dumpsys battery | grep temperature
33.9 °C. Not thermal.
The culprit is schedutil, Android's default governor. It scales frequency based on utilization estimates — and LLM inference is bursty: small work units, back to back. The estimate lags, the frequency hunts, and you get exactly this oscillation.
for c in 0 1 2 3 4 5 6 7; do
g=/sys/devices/system/cpu/cpu$c/cpufreq/scaling_governor
chmod 666 "$g"
echo performance > "$g"
done
Confirmed pinned: cpu0-3 @1785, cpu4-6 @2419, cpu7 @2841 MHz.
| Metric | schedutil | performance | Delta |
|---|---|---|---|
| Generation | 0.5 tok/s | 0.6 tok/s | +20% |
| First token | 1863 ms | 1619 ms | −13% |
20%. Throttling is real, but it is not the dominant factor.
| Hypothesis | Verdict |
|---|---|
Missing ARM optimization ( -march ) |
✅ Real, fixed — 2.5x on prefill |
| Threads scheduled onto little cores | ❌ Ruled out |
| OpenMP not actually threading | ❌ Ruled out |
| mmap pages evicted | ❌ Ruled out ( LLAMA_LOAD_MODE_MLOCK changed nothing) |
| CPU frequency throttling | ⚠️ Real, but only ~20% |
CPU at full clock (2419/2841 MHz)
threads on big cores (cpu4-7)
dotprod instructions present (verified by disassembly)
|
v
Utilization of peak compute capacity: ~1%
The arithmetic:
per token: 1.5B params x 2 = 3 GFLOPs
measured: 1.67 s per token
effective: 1.8 GFLOPS
peak: ~182 GFLOPS (4x A76 at full clock)
utilization: 1%
And memory:
per token reads all weights = 1.06 GB
1.06 GB / 1.67 s = 635 MB/s
available bandwidth: 17-34 GB/s
utilization: 2-4%
Neither the CPU nor the memory bus is anywhere near saturated. It's just slow.
The remaining explanation points inside ggml's quantized kernels on this platform — memory access patterns, cache behaviour, or OpenMP barrier overhead on small work units.
I can't fix that from configuration. It stays on the list as the next thing to dig into.
The investigation produced no final answer — but the process is reusable.
awk '{print $39}' /proc/<pid>/task/<tid>/stat # current CPU per thread
awk '{print $14+$15}' /proc/<pid>/task/<tid>/stat # cumulative CPU time per thread
cat /sys/devices/system/cpu/cpu*/cpufreq/scaling_cur_freq # live frequency
cat /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor # governor
Works on any Android device. No extra tooling.
Ruling out A, B and C is progress. It shrinks the search space from "anything" to "inside ggml's kernels".
If I only published the successful fix, you'd think this was a clean path. Real engineering is mostly wrong hypotheses and the value is eliminating them quickly.
Most people say "my model is slow" without quantifying it. Work out the memory-bandwidth ceiling and the compute ceiling, then measure. Knowing the gap is 60-120x tells you the problem isn't "needs tuning" — it's "something isn't running at all".
| Item | Restored to |
|---|---|
| CPU governor | schedutil |
| SELinux | Enforcing |
| stay-awake flag | false |
| screen timeout | 60000 ms |
| temp scripts | deleted |
Don't leave side effects on a test device — especially SELinux.
Next up: ggml quantized kernel efficiency on ARM.
If you've dug into this, I'd love to hear it.
Part 1: Why Your Phone Runs LLMs 80x Slower Than It Should
Part 2: this post
Code: https://github.com/Pingredsai/local-ai-android