cd /news/large-language-models/janus-run-local-llms-from-a-single-g… · home › topics › large-language-models › article
[ARTICLE · art-145148] src=byteiota.com ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

Janus: Run Local LLMs From a Single Go Binary — No Python, No Docker

Vibra-Ingenn released Janus, an open-source local LLM server written in Go that wraps llama.cpp inference behind an OpenAI-compatible REST API and ships as a single static binary with zero runtime dependencies. Janus serves on 127.0.0.1:8990, supports Vulkan GPU acceleration for AMD, Intel and NVIDIA plus CPU fallback, requires Go 1.22 or higher to build from source, and offers hot-swapping models via POST /models/load without a restart. The project positions itself against Ollama, whose roughly 680MB binary runs as a persistent daemon and has 166k GitHub stars, by targeting embedding, CI/CD scripting and air-gapped environments where Python or Docker are unavailable.

read4 min views5 publishedOct 5, 2026
Janus: Run Local LLMs From a Single Go Binary — No Python, No Docker
Image: Byteiota (auto-discovered)

Local LLMs are genuinely useful — until you try to actually set one up. The standard path means installing Python, creating a virtual environment, pulling in a stack of dependencies, maybe spinning up Docker, and keeping Ollama running as a background daemon that consumes RAM you didn’t budget for. All before writing a single line of application code. Janus is a new open-source Go project that cuts through every one of those steps. One binary. Drop a GGUF model file next to it. Get an OpenAI-compatible API on port 8990. That’s the whole setup.

What Janus Actually Is #

Janus is a local LLM server built in Go. It wraps llama.cpp inference behind an OpenAI-compatible REST API and ships as a single static binary with zero runtime dependencies. No Python interpreter, no Docker socket, no background service to manage. When you’re done, you stop the process. Nothing lingers.

GPU support comes through Vulkan — which means AMD, Intel, and NVIDIA all work out of the box. If you don’t have a supported GPU, it falls back to CPU. The project targets Windows 10/11, Linux, and macOS, and requires Go 1.22 or higher to build from source (prebuilt binaries are available).

Key features at a glance:

  • Full /v1/chat/completions and/v1/models API (OpenAI-compatible)
  • Streaming responses via SSE
  • Hot-swap models without restarting — POST /models/load handles it
  • Thinking model support with reasoning split into reasoning_content
  • Chat template auto-detection from GGUF metadata

Getting Started in Under Two Minutes #

Download the binary for your platform, put a GGUF model file in the same directory, and start Janus pointing at it:

./janus --model ./llama-3.2-3b.gguf

The server starts on 127.0.0.1:8990. From there, any OpenAI client works without modification:

curl http://127.0.0.1:8990/v1/chat/completions   -H "Content-Type: application/json"   -d '{
    "model": "local",
    "messages": [{"role": "user", "content": "Explain Go interfaces in one paragraph"}]
  }'

Point Cursor, Cline, or any tool that speaks the OpenAI API directly at this endpoint. Swap models mid-session without a restart using POST /models/load, or list what’s available with GET /models/list. It handles streaming, thinking model responses, and chat templates automatically.

Where Janus Beats Ollama (and Where It Doesn’t) #

Ollama is the obvious comparison, and it deserves credit — it’s the easiest way to get a local model running for personal use, and its 166k GitHub stars make it the largest open-source AI project by a significant margin. But Ollama ships as a ~680MB binary that runs as a persistent daemon. For many workflows, that’s fine. For others, it’s friction you don’t need.

Scenario Janus Ollama
Zero-install project embed Win Requires daemon
CI/CD pipeline scripting Win Overkill
Air-gapped / secure environments Win Workable
Cross-GPU (AMD / Intel) Win (Vulkan) Mixed
Model library and discovery Bring your own Win
GUI desktop experience Not applicable Win (Open WebUI)

Janus is a runtime, not a platform. It doesn’t manage a model library or provide a GUI — you bring the GGUF files, it serves them. If you need ollama pull llama3.2 and a browser UI, Janus isn’t the right tool. If you need to ship local inference as part of an application, embed it in a CI test harness, or run it in an environment where Python is unavailable, Janus removes every blocker.

The Go Angle Is Not a Coincidence #

Janus, Ollama, goinfer — three of the most practical local LLM tools share the same language. Go’s static compilation is uniquely suited to this class of problem. Single binary. No interpreter. No shared library hell. Deploy by copying a file.

Python still dominates ML training and research, and that’s unlikely to change. But Go is quietly taking the “serve and ship” layer of AI tooling the same way it took CLI tools from Python and Ruby between 2015 and 2020. Caddy, kubectl, Terraform — the pattern is familiar. The zero-dependency distribution model wins when the audience is developers integrating tools rather than researchers running experiments.

Vulkan support accelerates this. CUDA works, but it locks you to NVIDIA. Vulkan runs across the GPU landscape, which matters more now that AMD and Intel have meaningful market share in developer machines. Tools that ship Vulkan-first have a structural advantage for widespread adoption.

What to Watch #

Janus is early-stage — it surfaced on Hacker News this week and is under active development. The core functionality is solid, but model management (compared to Ollama’s pull-and-run workflow) requires sourcing GGUF files yourself from Hugging Face or similar. That’s a real gap for developers new to local LLMs.

What makes it worth watching now: the philosophy is correct, the implementation is clean, and the timing is right. Developer appetite for zero-dependency local AI tooling is only growing. Janus is available on GitHub today. The Show HN thread has early community feedback worth reading before you adopt it in a production workflow.

── more in #large-language-models 4 stories · sorted by recency
── more on @janus 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/janus-run-local-llms…] indexed:0 read:4min 2026-10-05 · —