{"slug": "localchat-yet-another-local-chat-app-poc-w-golang-htmx", "title": "localchat: yet another local chat app POC w/GoLang + HTMX", "summary": "A developer built localchat, a proof-of-concept local AI chat app pairing a Go 1.27 server with Google's Gemma 4 E2B model running on the user's own machine. The app streams tokens into the browser as plain HTML over server-sent events using HTMX 2's SSE extension, with no JavaScript written for the streaming, and keeps conversations in memory per browser session. On an M3 MacBook Air, the 4-bit Gemma 4 E2B emits its first token in about a second and finishes a short answer with a code block in 3 to 5 seconds.", "body_md": "*This is a submission for the [Hacktoberfest Weekend Challenge: Build for a Friend](https://dev.to/challenges/hacktoberfest-weekend-2026-10-01)*\n\nlocalchat is a chat app for an AI model that runs on your own machine. A Go server renders the UI, and Google's **Gemma 4 E2B** writes the replies, streamed into your browser token by token. Your messages stay in memory on your laptop for as long as you run the app.\n\nI talked to a friend before about what a bare-basic and simple Go web app with AI integration would look like, without too many difficult parts. So I spent some time setting this up and kept the project small enough to read in one sitting.\n\nI picked Gemma while developing because I like running small models on my own hardware. Gemma 4 E2B fits in a few gigabytes of RAM and starts answering in about a second on my laptop. It's fast and easy to develop against while still having realistic output.\n\nMy friend and I get two things:\n\n**Repo:** [git.b0b.be/bdeb/localchat](https://git.b0b.be/bdeb/localchat)\n\n```\ncmd/localchat/          entry point: serve + install/start/stop as an OS service\ninternal/llm/           ~150-line OpenAI-compatible streaming client (stdlib only)\ninternal/chat/          in-memory conversations, one per browser session\ninternal/web/           routes, SSE streaming, embedded static assets\ninternal/web/views/     Templ components\nscripts/e2e.py          Playwright browser test against the real model\ncompose.yaml            app + llama.cpp + Gemma 4 E2B\n```\n\nOn my M3 MacBook Air, Gemma 4 E2B (4-bit, through oMLX) sends its first token after about a second and finishes a short answer with a code block in 3 to 5 seconds.\n\n**Open-source AI:** Google's Gemma 4 E2B (open weights, 4-bit). I ran it with **oMLX** (Apple MLX) on my Mac while developing, and the container uses **llama.cpp** (`llama-server` with a GGUF build).\n\n**App stack:** Go 1.27, Templ, HTMX 2 with its SSE extension, goldmark and kardianos/service.\n\n```\nBrowser ── HTMX + SSE extension\n  │  POST /chat              → user bubble + empty reply bubble (sse-connect)\n  │  GET  /chat/stream/{id}  ← \"token\" events (append) … \"done\" (swap in Markdown)\nGo (net/http + Templ)\n  │  POST /v1/chat/completions {stream: true}\nLocal model server ── Gemma 4 E2B\n```\n\nStreaming works as plain HTML over server-sent events:\n\n`sse-connect=\"/chat/stream/{id}\"`.`<span>` inside a `token` event. HTMX appends each span with `hx-swap=\"beforeend\"`, so you see the answer appear word by word. I wrote no JavaScript for the streaming.`done` event carrying the reply as server-rendered Markdown (goldmark, raw HTML stripped). HTMX swaps the whole bubble for it, and removing the `sse-connect` element closes the stream.\nA few more details:\n\n`bufio.Scanner` over `data:` lines, in about 150 lines of standard-library Go.`localchat install --user` registers the app as a launchd agent, a systemd unit or a Windows service, pointed at your `.env` file.", "url": "https://wpnews.pro/news/localchat-yet-another-local-chat-app-poc-w-golang-htmx", "canonical_source": "https://dev.to/bdeb1337/localchat-yet-another-local-chat-app-poc-wgolang-htmx-265d", "published_at": "2026-10-05 06:01:14+00:00", "updated_at": "2026-10-05 06:13:03.609714+00:00", "lang": "en", "topics": ["ai-tools", "large-language-models", "developer-tools", "ai-products"], "entities": ["localchat", "Go", "HTMX", "Gemma 4 E2B", "Google", "oMLX", "llama.cpp", "Templ"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/localchat-yet-another-local-chat-app-poc-w-golang-htmx", "markdown": "https://wpnews.pro/news/localchat-yet-another-local-chat-app-poc-w-golang-htmx.md", "text": "https://wpnews.pro/news/localchat-yet-another-local-chat-app-poc-w-golang-htmx.txt", "jsonld": "https://wpnews.pro/news/localchat-yet-another-local-chat-app-poc-w-golang-htmx.jsonld"}}