{"slug": "janus-run-local-llms-from-a-single-go-binary-no-python-no-docker", "title": "Janus: Run Local LLMs From a Single Go Binary — No Python, No Docker", "summary": "Vibra-Ingenn released Janus, an open-source local LLM server written in Go that wraps llama.cpp inference behind an OpenAI-compatible REST API and ships as a single static binary with zero runtime dependencies. Janus serves on 127.0.0.1:8990, supports Vulkan GPU acceleration for AMD, Intel and NVIDIA plus CPU fallback, requires Go 1.22 or higher to build from source, and offers hot-swapping models via POST /models/load without a restart. The project positions itself against Ollama, whose roughly 680MB binary runs as a persistent daemon and has 166k GitHub stars, by targeting embedding, CI/CD scripting and air-gapped environments where Python or Docker are unavailable.", "body_md": "Local LLMs are genuinely useful — until you try to actually set one up. The standard path means installing Python, creating a virtual environment, pulling in a stack of dependencies, maybe spinning up Docker, and keeping Ollama running as a background daemon that consumes RAM you didn’t budget for. All before writing a single line of application code. [Janus](https://github.com/Vibra-Ingenn/Janus) is a new open-source Go project that cuts through every one of those steps. One binary. Drop a GGUF model file next to it. Get an OpenAI-compatible API on port 8990. That’s the whole setup.\n\n## What Janus Actually Is\n\nJanus is a local LLM server built in Go. It wraps [llama.cpp](https://github.com/ggml-org/llama.cpp) inference behind an OpenAI-compatible REST API and ships as a single static binary with zero runtime dependencies. No Python interpreter, no Docker socket, no background service to manage. When you’re done, you stop the process. Nothing lingers.\n\nGPU support comes through Vulkan — which means AMD, Intel, and NVIDIA all work out of the box. If you don’t have a supported GPU, it falls back to CPU. The project targets Windows 10/11, Linux, and macOS, and requires Go 1.22 or higher to build from source (prebuilt binaries are available).\n\nKey features at a glance:\n\n- Full `/v1/chat/completions` and`/v1/models` API (OpenAI-compatible)\n- Streaming responses via SSE\n- Hot-swap models without restarting — `POST /models/load` handles it\n- Thinking model support with reasoning split into `reasoning_content`\n- Chat template auto-detection from GGUF metadata\n\n## Getting Started in Under Two Minutes\n\nDownload the binary for your platform, put a GGUF model file in the same directory, and start Janus pointing at it:\n\n```\n./janus --model ./llama-3.2-3b.gguf\n```\n\nThe server starts on `127.0.0.1:8990`. From there, any OpenAI client works without modification:\n\n```\ncurl http://127.0.0.1:8990/v1/chat/completions   -H \"Content-Type: application/json\"   -d '{\n    \"model\": \"local\",\n    \"messages\": [{\"role\": \"user\", \"content\": \"Explain Go interfaces in one paragraph\"}]\n  }'\n```\n\nPoint Cursor, Cline, or any tool that speaks the OpenAI API directly at this endpoint. Swap models mid-session without a restart using `POST /models/load`, or list what’s available with `GET /models/list`. It handles streaming, thinking model responses, and chat templates automatically.\n\n## Where Janus Beats Ollama (and Where It Doesn’t)\n\nOllama is the obvious comparison, and it deserves credit — it’s the easiest way to get a local model running for personal use, and its [166k GitHub stars](https://github.com/ollama/ollama) make it the largest open-source AI project by a significant margin. But Ollama ships as a ~680MB binary that runs as a persistent daemon. For many workflows, that’s fine. For others, it’s friction you don’t need.\n\n| Scenario | Janus | Ollama | \n|---|---|---|\n| Zero-install project embed | Win | Requires daemon | \n| CI/CD pipeline scripting | Win | Overkill | \n| Air-gapped / secure environments | Win | Workable | \n| Cross-GPU (AMD / Intel) | Win (Vulkan) | Mixed | \n| Model library and discovery | Bring your own | Win | \n| GUI desktop experience | Not applicable | Win (Open WebUI) | \n\nJanus is a runtime, not a platform. It doesn’t manage a model library or provide a GUI — you bring the GGUF files, it serves them. If you need `ollama pull llama3.2` and a browser UI, Janus isn’t the right tool. If you need to ship local inference as part of an application, embed it in a CI test harness, or run it in an environment where Python is unavailable, Janus removes every blocker.\n\n## The Go Angle Is Not a Coincidence\n\nJanus, Ollama, goinfer — three of the most practical local LLM tools share the same language. Go’s static compilation is uniquely suited to this class of problem. Single binary. No interpreter. No shared library hell. Deploy by copying a file.\n\nPython still dominates ML training and research, and that’s unlikely to change. But Go is quietly taking the “serve and ship” layer of AI tooling the same way it took CLI tools from Python and Ruby between 2015 and 2020. Caddy, kubectl, Terraform — the pattern is familiar. The zero-dependency distribution model wins when the audience is developers integrating tools rather than researchers running experiments.\n\nVulkan support accelerates this. CUDA works, but it locks you to NVIDIA. Vulkan runs across the GPU landscape, which matters more now that AMD and Intel have meaningful market share in developer machines. Tools that ship Vulkan-first have a structural advantage for widespread adoption.\n\n## What to Watch\n\nJanus is early-stage — it surfaced on Hacker News this week and is under active development. The core functionality is solid, but model management (compared to Ollama’s pull-and-run workflow) requires sourcing GGUF files yourself from [Hugging Face](https://huggingface.co/models?library=gguf) or similar. That’s a real gap for developers new to local LLMs.\n\nWhat makes it worth watching now: the philosophy is correct, the implementation is clean, and the timing is right. Developer appetite for zero-dependency local AI tooling is only growing. Janus is [available on GitHub](https://github.com/Vibra-Ingenn/Janus) today. The [Show HN thread](https://news.ycombinator.com/item?id=49926773) has early community feedback worth reading before you adopt it in a production workflow.", "url": "https://wpnews.pro/news/janus-run-local-llms-from-a-single-go-binary-no-python-no-docker", "canonical_source": "https://byteiota.com/janus-run-local-llms-from-a-single-go-binary-no-python-no-docker/", "published_at": "2026-10-05 04:08:15+00:00", "updated_at": "2026-10-05 04:12:54.639739+00:00", "lang": "en", "topics": ["large-language-models", "ai-tools", "developer-tools", "ai-infrastructure"], "entities": ["Janus", "Vibra-Ingenn", "llama.cpp", "Ollama", "Go", "Vulkan", "OpenAI", "Cursor"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/janus-run-local-llms-from-a-single-go-binary-no-python-no-docker", "markdown": "https://wpnews.pro/news/janus-run-local-llms-from-a-single-go-binary-no-python-no-docker.md", "text": "https://wpnews.pro/news/janus-run-local-llms-from-a-single-go-binary-no-python-no-docker.txt", "jsonld": "https://wpnews.pro/news/janus-run-local-llms-from-a-single-go-binary-no-python-no-docker.jsonld"}}