localchat: yet another local chat app POC w/GoLang + HTMX A developer built localchat, a proof-of-concept local AI chat app pairing a Go 1.27 server with Google's Gemma 4 E2B model running on the user's own machine. The app streams tokens into the browser as plain HTML over server-sent events using HTMX 2's SSE extension, with no JavaScript written for the streaming, and keeps conversations in memory per browser session. On an M3 MacBook Air, the 4-bit Gemma 4 E2B emits its first token in about a second and finishes a short answer with a code block in 3 to 5 seconds. This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend https://dev.to/challenges/hacktoberfest-weekend-2026-10-01 localchat is a chat app for an AI model that runs on your own machine. A Go server renders the UI, and Google's Gemma 4 E2B writes the replies, streamed into your browser token by token. Your messages stay in memory on your laptop for as long as you run the app. I talked to a friend before about what a bare-basic and simple Go web app with AI integration would look like, without too many difficult parts. So I spent some time setting this up and kept the project small enough to read in one sitting. I picked Gemma while developing because I like running small models on my own hardware. Gemma 4 E2B fits in a few gigabytes of RAM and starts answering in about a second on my laptop. It's fast and easy to develop against while still having realistic output. My friend and I get two things: Repo: git.b0b.be/bdeb/localchat https://git.b0b.be/bdeb/localchat cmd/localchat/ entry point: serve + install/start/stop as an OS service internal/llm/ ~150-line OpenAI-compatible streaming client stdlib only internal/chat/ in-memory conversations, one per browser session internal/web/ routes, SSE streaming, embedded static assets internal/web/views/ Templ components scripts/e2e.py Playwright browser test against the real model compose.yaml app + llama.cpp + Gemma 4 E2B On my M3 MacBook Air, Gemma 4 E2B 4-bit, through oMLX sends its first token after about a second and finishes a short answer with a code block in 3 to 5 seconds. Open-source AI: Google's Gemma 4 E2B open weights, 4-bit . I ran it with oMLX Apple MLX on my Mac while developing, and the container uses llama.cpp llama-server with a GGUF build . App stack: Go 1.27, Templ, HTMX 2 with its SSE extension, goldmark and kardianos/service. Browser ── HTMX + SSE extension │ POST /chat → user bubble + empty reply bubble sse-connect │ GET /chat/stream/{id} ← "token" events append … "done" swap in Markdown Go net/http + Templ │ POST /v1/chat/completions {stream: true} Local model server ── Gemma 4 E2B Streaming works as plain HTML over server-sent events: sse-connect="/chat/stream/{id}" .