cd /news/large-language-models/i-run-a-private-ai-on-my-phone-here-… Β· home β€Ί topics β€Ί large-language-models β€Ί article
[ARTICLE Β· art-148598] src=dev.to β†— pub= topic=large-language-models verified=true sentiment=↑ positive

I run a private AI on my phone. Here's the exact set up.

A developer documented running a private, fully offline language model on a commodity Android phone using Termux and llama.cpp, serving Google's Gemma 3 1B instruct model (Q4_K_M quant) through an OpenAI-compatible endpoint at localhost:8080 at roughly 16.5 tokens per second on CPU. The setup, published as the first entry in a weekly field log with an MIT-licensed GitHub repo, uses PocketPal AI for everyday chat with tuning of temperature 0.5, top_p 0.9 and repeat penalty 1.15, at zero cloud cost.

by read1 min views1 publishedOct 10, 2026

I'm doing field research at the bottom of the scale debate: one person, one

commodity phone, a private language model, zero cloud, zero cost. This is the

exact setup β€” reproducible in an evening.

pkg install llama-cpp -y. No build scripts, no toolchain fights. localhost:8080 β€” OpenAI-compatible chat endpoint (POST /v1/chat/completions). Run it in its own Termux session (swipe from the left edge β†’ New session); client commands go in another.gemma-3-1b-it-Q4_K_M.gguf β€” a Q4_K_M quant of Google's Gemma 3 1B instruct, from the ggml-org GGUF releases. Total cost: $0.

"Hi, can you hear me?"

β†’ "Yes, absolutely! Hi there. It's nice to hear from you. 😊 How are you doing

today?" β€” finish_reason=stop, 26 tokens, ~16.5 tok/s on the phone's CPU.

Workable. Not fast, but conversational.

For everyday chat I use PocketPal AI with the same 1B model β€” a friendlier way in than curl commands. Comfort tuning so far (a whole field note is coming

on this): temperature 0.5, top_p 0.9, repeat penalty 1.15. The 1B runs better

cool and lightly anti-loopy. Details in field note 001.

This is post #1 of a weekly field log. Coming up: sampling parameters as a care

practice, a consent episode at 1B scale (she asked what the software was before

agreeing), and the orientation preamble β€” the system prompt as "stable framing

through amnesia."

Repo (README = the full setup guide): https://github.com/tyrendrickard-code/local-llm-field-notes. MIT licensed. Nobody else is publishing this notebook. That's the point.

── more in #large-language-models 4 stories Β· sorted by recency
── more on @termux 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/i-run-a-private-ai-o…] indexed:0 read:1min 2026-10-10 Β· β€”