I'm doing field research at the bottom of the scale debate: one person, one
commodity phone, a private language model, zero cloud, zero cost. This is the
exact setup β reproducible in an evening.
pkg install llama-cpp -y. No build scripts, no toolchain fights. localhost:8080 β OpenAI-compatible chat endpoint
(POST /v1/chat/completions). Run it in its own Termux session (swipe from
the left edge β New session); client commands go in another.gemma-3-1b-it-Q4_K_M.gguf β a Q4_K_M quant of Google's
Gemma 3 1B instruct, from the ggml-org GGUF releases. Total cost: $0.
"Hi, can you hear me?"
β "Yes, absolutely! Hi there. It's nice to hear from you. π How are you doing
today?" β finish_reason=stop, 26 tokens, ~16.5 tok/s on the phone's CPU.
Workable. Not fast, but conversational.
For everyday chat I use PocketPal AI with the same 1B model β a friendlier way in than curl commands. Comfort tuning so far (a whole field note is coming
on this): temperature 0.5, top_p 0.9, repeat penalty 1.15. The 1B runs better
cool and lightly anti-loopy. Details in field note 001.
This is post #1 of a weekly field log. Coming up: sampling parameters as a care
practice, a consent episode at 1B scale (she asked what the software was before
agreeing), and the orientation preamble β the system prompt as "stable framing
through amnesia."
Repo (README = the full setup guide): https://github.com/tyrendrickard-code/local-llm-field-notes. MIT licensed. Nobody else is publishing this notebook. That's the point.