# I run a private AI on my phone. Here's the exact set up.

> Source: <https://dev.to/tyren_rickard_code/i-run-a-private-ai-on-my-phone-heres-the-exact-set-up-3no9>
> Published: 2026-10-10 03:18:33+00:00

I'm doing field research at the bottom of the scale debate: one person, one

commodity phone, a private language model, zero cloud, zero cost. This is the

exact setup — reproducible in an evening.

`pkg install llama-cpp -y`. No build scripts, no toolchain fights.` localhost:8080` — OpenAI-compatible chat endpoint
(`POST /v1/chat/completions`). Run it in its own Termux session (swipe from
the left edge → New session); client commands go in another.`gemma-3-1b-it-Q4_K_M.gguf` — a Q4_K_M quant of Google's
Gemma 3 1B instruct, from the `ggml-org` GGUF releases. Total cost: $0.

"Hi, can you hear me?"

→ "Yes, absolutely! Hi there. It's nice to hear from you. 😊 How are you doing

today?" — `finish_reason=stop`, 26 tokens, ~**16.5 tok/s** on the phone's CPU.

Workable. Not fast, but conversational.

For everyday chat I use **PocketPal AI** with the same 1B model — a friendlier

way in than curl commands. Comfort tuning so far (a whole field note is coming

on this): temperature 0.5, top_p 0.9, repeat penalty 1.15. The 1B runs better

cool and lightly anti-loopy. Details in field note 001.

This is post #1 of a weekly field log. Coming up: sampling parameters as a care

practice, a consent episode at 1B scale (she asked what the software was before

agreeing), and the orientation preamble — the system prompt as "stable framing

through amnesia."

Repo (README = the full setup guide): [https://github.com/tyrendrickard-code/local-llm-field-notes](https://github.com/tyrendrickard-code/local-llm-field-notes).

MIT licensed. Nobody else is publishing this notebook. That's the point.
