# On-Device Inference Debugging (Part 2): Threads on Big Cores, CPU at Full Clock — Still Slow

> Source: <https://dev.to/pingredsai/on-device-inference-debugging-part-2-threads-on-big-cores-cpu-at-full-clock-still-slow-6hm>
> Published: 2026-10-10 11:02:50+00:00

Part 1: *Why Your Phone Runs LLMs 80x Slower Than It Should*

This part focuses on **thread scheduling** and **CPU frequency scaling**.

In part 1 I benchmarked Qwen2.5-1.5B on a Pixel 4 (Snapdragon 855) and measured **0.5 tok/s** — 60-120x off the theoretical limits for that SoC.

That investigation found one real bug: **llama.cpp cross-compilation passes no `-march` by default**, so quantized matmul falls back to a slow path. Fixing it cut prefill from 114 s to 45 s.

**But throughput was still 0.5 tok/s.** The biggest hole was filled; the problem remained.

After publishing on DEV, a reader offered two traps that *look* like a slow model but aren't:

**Scheduling:** a llama.cpp server whose child processes were parked on efficiency cores under load — a 3-second transcription became 37 seconds. Same binary, same flags; the only difference was who launched it.

**Storage:** running off a USB stick, the whole model is re-read at every start — 2.6 GB on a 34 MB/s drive means over a minute before the first token.

And the line that stung:

**The number that looks like "the model is slow" is almost never the model.**

Both point the same way: maybe the problem isn't *what* is being computed, but *where* and *how fast*.

```
for c in /sys/devices/system/cpu/cpu[0-9]*; do
  echo "$(basename $c) $(cat $c/cpufreq/cpuinfo_max_freq)"
done
```

| Cores | Max freq | Type | 
|---|---|---|
| cpu0–3 | 1.79 GHz | A55 little | 
| cpu4–6 | 2.42 GHz | A76 mid | 
| cpu7 | **2.84 GHz** | A76 big | 

Classic 1+3+4. If inference threads land on cpu0–3, you lose 3-4x immediately.

Field 39 of `/proc/<pid>/task/<tid>/stat` is **the CPU the thread is currently running on**.

```
pid=$(pidof com.example.localai)

for t in /proc/$pid/task/*; do
  tid=$(basename $t)
  comm=$(cat $t/comm)
  cpu=$(awk '{print $39}' $t/stat)
  ticks=$(awk '{print $14+$15}' $t/stat)
  echo "$tid $comm cpu=$cpu ticks=$ticks"
done
```

**Sample repeatedly** — a single snapshot tells you nothing.

``` php
main thread 9556:  cpu5 -> cpu5 -> cpu5 -> cpu3 -> cpu5 -> cpu4
19524 openmp_worker  cpu6
19525 openmp_worker  cpu7   <- big core
19526 openmp_worker  cpu5
```

**All four inference threads stayed on cpu4–7. None touched the little cores.**

Also confirmed OpenMP is genuinely working:

```
19524 openmp_worker   5607 ticks
19525 openmp_worker   5606 ticks
19526 openmp_worker   5602 ticks
9556  DefaultDispatch 12132 ticks (main)
```

`n_threads=4` → 3 workers + 1 main thread. Exactly as configured.

**Verdict: "threads parked on efficiency cores" does not happen on this device. Ruled out.**

```
cpu4:  710 ~ 1920 MHz   (max 2419)   <- as low as 29%
cpu5: 1612 ~ 2419 MHz   (max 2419)
cpu6:  710 ~ 2419 MHz   (max 2419)
cpu7: 1612 ~ 2016 MHz   (max 2841)   <- never above 71%
php
adb shell dumpsys battery | grep temperature
# temperature: 339  ->  33.9 C
```

33.9 °C. Not thermal.

The culprit is **`schedutil`**, Android's default governor. It scales frequency based on utilization estimates — and LLM inference is *bursty*: small work units, back to back. The estimate lags, the frequency hunts, and you get exactly this oscillation.

```
# root required; SELinux Enforcing blocks a direct write, so chmod first
for c in 0 1 2 3 4 5 6 7; do
  g=/sys/devices/system/cpu/cpu$c/cpufreq/scaling_governor
  chmod 666 "$g"
  echo performance > "$g"
done
```

Confirmed pinned: cpu0-3 @1785, cpu4-6 @2419, cpu7 @2841 MHz.

| Metric | schedutil | performance | Delta | 
|---|---|---|---|
| Generation | 0.5 tok/s | **0.6 tok/s** | **+20%** | 
| First token | 1863 ms | 1619 ms | −13% | 

**20%.** Throttling is real, but it is not the dominant factor.

| Hypothesis | Verdict | 
|---|---|
| Missing ARM optimization ( `-march` ) | ✅ **Real, fixed** — 2.5x on prefill | 
| Threads scheduled onto little cores | ❌ Ruled out | 
| OpenMP not actually threading | ❌ Ruled out | 
| mmap pages evicted | ❌ Ruled out ( `LLAMA_LOAD_MODE_MLOCK` changed nothing) | 
| CPU frequency throttling | ⚠️ Real, but only ~20% | 

```
CPU at full clock (2419/2841 MHz)
threads on big cores (cpu4-7)
dotprod instructions present (verified by disassembly)
        |
        v
Utilization of peak compute capacity: ~1%
```

The arithmetic:

```
per token: 1.5B params x 2 = 3 GFLOPs
measured: 1.67 s per token
effective: 1.8 GFLOPS
peak: ~182 GFLOPS (4x A76 at full clock)
utilization: 1%
```

And memory:

```
per token reads all weights = 1.06 GB
1.06 GB / 1.67 s = 635 MB/s
available bandwidth: 17-34 GB/s
utilization: 2-4%
```

**Neither the CPU nor the memory bus is anywhere near saturated. It's just slow.**

The remaining explanation points inside **ggml's quantized kernels on this platform** — memory access patterns, cache behaviour, or OpenMP barrier overhead on small work units.

**I can't fix that from configuration. It stays on the list as the next thing to dig into.**

The investigation produced no final answer — but the *process* is reusable.

```
awk '{print $39}' /proc/<pid>/task/<tid>/stat      # current CPU per thread
awk '{print $14+$15}' /proc/<pid>/task/<tid>/stat  # cumulative CPU time per thread
cat /sys/devices/system/cpu/cpu*/cpufreq/scaling_cur_freq   # live frequency
cat /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor   # governor
```

**Works on any Android device. No extra tooling.**

Ruling out A, B and C *is* progress. It shrinks the search space from "anything" to "inside ggml's kernels".

If I only published the successful fix, you'd think this was a clean path. Real engineering is mostly wrong hypotheses and the value is eliminating them quickly.

Most people say "my model is slow" without quantifying it. Work out the memory-bandwidth ceiling and the compute ceiling, then measure. **Knowing the gap is 60-120x tells you the problem isn't "needs tuning" — it's "something isn't running at all".**

| Item | Restored to | 
|---|---|
| CPU governor | `schedutil` | 
| SELinux | `Enforcing` | 
| stay-awake flag | `false` | 
| screen timeout | 60000 ms | 
| temp scripts | deleted | 

**Don't leave side effects on a test device — especially SELinux.**

Next up: **ggml quantized kernel efficiency on ARM**.

If you've dug into this, I'd love to hear it.

*Part 1: [Why Your Phone Runs LLMs 80x Slower Than It Should](https://dev.to/pingredsai/why-your-phone-runs-llms-80x-slower-than-it-should-and-what-i-found-2645)*

*Part 2: this post*

*Code: [https://github.com/Pingredsai/local-ai-android](https://github.com/Pingredsai/local-ai-android)*
