Janus: Run GGUF Models Locally With One Go Binary Janus, an MIT-licensed single Go binary from the Vibra-Ingenn project, launched as a Show HN on October 1 with 42+ points, wrapping llama.cpp's Vulkan backend to run GGUF models on AMD, Intel, or NVIDIA GPUs with CPU fallback and no Python, Docker, or Ollama daemon. Janus exposes an OpenAI-compatible API on localhost via /v1/chat/completions and /v1/models, adds hot-swap model loading through POST /models/load and OpenRouter cloud fallback via INFERENCE_BACKEND=openrouter, and requires Go 1.22+ on Linux while macOS is limited to CPU-only inference. Hacker News commenters noted Intel GPU Vulkan overhead can undercut CPU-only gains depending on driver version, and no published benchmarks compare Janus to vLLM or SGLang. Setting up local LLM inference has always carried a tax: install Python, configure a virtual environment, wrestle with CUDA drivers, debug ROCm compatibility, then finally load your model. Janus https://github.com/Vibra-Ingenn/Janus eliminates every step between “download binary” and “query your model.” It’s a single MIT-licensed Go binary that wraps llama.cpp’s Vulkan backend, exposes an OpenAI-compatible API on localhost, and runs GGUF models on AMD, Intel, or NVIDIA GPUs with CPU fallback — no Python, no Docker, no Ollama daemon required. The project surfaced as a Show HN on October 1 with 42+ points. What Janus Is Janus is not a service, a registry, or a managed runtime. It’s one executable. The binary ships with pre-compiled llama.cpp Vulkan DLLs on Windows, so there’s no C++ build step. Drop a .gguf file in the models/ directory, point the JANUS MODEL PATH environment variable at it, and the server is up. The API surface is intentionally minimal: /v1/chat/completions and /v1/models follow the OpenAI spec exactly, which means any OpenAI-compatible client works without modification. Beyond standard inference, Janus adds two features worth noting. First, hot-swap: you can change the active model without restarting the server by calling POST /models/load with a new path. Second, cloud fallback: set INFERENCE BACKEND=openrouter to route requests to OpenRouter https://openrouter.ai when your local GPU is saturated or you want to switch to a larger model. That’s a genuinely useful escape hatch for development workflows where you prototype locally but occasionally need a heavier model. Getting Started The build process is minimal. On Windows, build.ps1 downloads the pre-built llama.cpp Vulkan DLLs and compiles dist\janus.exe in one step. Linux requires Go 1.22+ and a Vulkan-capable driver. macOS users get CPU-only inference — Vulkan support on macOS is not guaranteed and Ollama remains the better choice there. Clone and build git clone https://github.com/Vibra-Ingenn/Janus cd Janus ./build.ps1 Windows make build Linux Drop your model in models/ cp ~/Downloads/llama3.2-3b-q4 k m.gguf models/ Run the server JANUS MODEL PATH=models/llama3.2-3b-q4 k m.gguf ./janus Query it — identical to OpenAI API calls curl http://localhost:8080/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{"model":"local","messages": {"role":"user","content":"Hello"} }' VRAM tuning uses two variables: JANUS GPU LAYERS -1 offloads all layers to GPU, 0 forces CPU and JANUS VRAM CEILING MB sets a budget cap in MiB. Hit /engine/status to see live VRAM usage and the active backend. Janus vs Ollama vs llama.cpp Server Ollama https://ollama.com has become the default local inference choice for most developers, and it deserves that position. Its built-in model registry, ollama pull workflow, and Metal acceleration on Apple Silicon are genuinely polished. Janus is not trying to replace that. The differences are narrow but meaningful: | | Janus | Ollama | llama.cpp server | |---|---|---|---| | Install | Single binary | Single binary + daemon | Build from C++ | | GPU backends | Vulkan | CUDA / ROCm / Metal | CUDA, ROCm, Metal, Vulkan | | Apple Silicon | CPU only | Metal fast | Metal fast | | Hot-swap models | Yes via API | No restart required | No | | Cloud fallback | OpenRouter | None | None | | Model registry | No raw .gguf | Built-in | No | | API | OpenAI-compat | OpenAI-compat | OpenAI-compat | The Hacker News thread flagged two real concerns. Intel GPU users report that Vulkan overhead can undercut performance gains versus CPU-only inference, depending on driver version. And there are no published benchmarks comparing Janus to vLLM https://docs.vllm.ai or SGLang — the target audience is single-user local inference, not production serving, but that should be stated explicitly. Where Janus Fits Three scenarios favor Janus over Ollama. One: CI/CD pipelines where dropping a binary and setting two environment variables is cleaner than running an Ollama daemon in a container. Two: Go projects where adding a Python runtime or a separate Ollama service creates real dependency friction. Three: mixed-GPU development teams where one developer has NVIDIA, another has AMD, and the Vulkan backend normalizes the difference without separate driver paths. For Apple Silicon users, Ollama’s Metal integration is faster and Janus adds nothing. For high-concurrency production use, neither tool applies — vLLM is the answer. And for developers who want a model registry with one-command downloads, Ollama’s workflow is smoother. Janus is a precision tool. Its value proposition is exactly what it says: one binary, no runtime dependencies, local GGUF inference, OpenAI-compatible API. For the workflows where that matters, it’s hard to beat.