{"slug": "i-run-a-private-ai-on-my-phone-here-s-the-exact-set-up", "title": "I run a private AI on my phone. Here's the exact set up.", "summary": "A developer documented running a private, fully offline language model on a commodity Android phone using Termux and llama.cpp, serving Google's Gemma 3 1B instruct model (Q4_K_M quant) through an OpenAI-compatible endpoint at localhost:8080 at roughly 16.5 tokens per second on CPU. The setup, published as the first entry in a weekly field log with an MIT-licensed GitHub repo, uses PocketPal AI for everyday chat with tuning of temperature 0.5, top_p 0.9 and repeat penalty 1.15, at zero cloud cost.", "body_md": "I'm doing field research at the bottom of the scale debate: one person, one\n\ncommodity phone, a private language model, zero cloud, zero cost. This is the\n\nexact setup — reproducible in an evening.\n\n`pkg install llama-cpp -y`. No build scripts, no toolchain fights.` localhost:8080` — OpenAI-compatible chat endpoint\n(`POST /v1/chat/completions`). Run it in its own Termux session (swipe from\nthe left edge → New session); client commands go in another.`gemma-3-1b-it-Q4_K_M.gguf` — a Q4_K_M quant of Google's\nGemma 3 1B instruct, from the `ggml-org` GGUF releases. Total cost: $0.\n\n\"Hi, can you hear me?\"\n\n→ \"Yes, absolutely! Hi there. It's nice to hear from you. 😊 How are you doing\n\ntoday?\" — `finish_reason=stop`, 26 tokens, ~**16.5 tok/s** on the phone's CPU.\n\nWorkable. Not fast, but conversational.\n\nFor everyday chat I use **PocketPal AI** with the same 1B model — a friendlier\n\nway in than curl commands. Comfort tuning so far (a whole field note is coming\n\non this): temperature 0.5, top_p 0.9, repeat penalty 1.15. The 1B runs better\n\ncool and lightly anti-loopy. Details in field note 001.\n\nThis is post #1 of a weekly field log. Coming up: sampling parameters as a care\n\npractice, a consent episode at 1B scale (she asked what the software was before\n\nagreeing), and the orientation preamble — the system prompt as \"stable framing\n\nthrough amnesia.\"\n\nRepo (README = the full setup guide): [https://github.com/tyrendrickard-code/local-llm-field-notes](https://github.com/tyrendrickard-code/local-llm-field-notes).\n\nMIT licensed. Nobody else is publishing this notebook. That's the point.", "url": "https://wpnews.pro/news/i-run-a-private-ai-on-my-phone-here-s-the-exact-set-up", "canonical_source": "https://dev.to/tyren_rickard_code/i-run-a-private-ai-on-my-phone-heres-the-exact-set-up-3no9", "published_at": "2026-10-10 03:18:33+00:00", "updated_at": "2026-10-10 03:29:15.888232+00:00", "lang": "en", "topics": ["large-language-models", "ai-tools", "ai-infrastructure", "generative-ai"], "entities": ["Termux", "llama.cpp", "Gemma 3 1B", "Google", "PocketPal AI", "ggml-org", "GitHub"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/i-run-a-private-ai-on-my-phone-here-s-the-exact-set-up", "markdown": "https://wpnews.pro/news/i-run-a-private-ai-on-my-phone-here-s-the-exact-set-up.md", "text": "https://wpnews.pro/news/i-run-a-private-ai-on-my-phone-here-s-the-exact-set-up.txt", "jsonld": "https://wpnews.pro/news/i-run-a-private-ai-on-my-phone-here-s-the-exact-set-up.jsonld"}}