{"slug": "janus-run-gguf-models-locally-with-one-go-binary", "title": "Janus: Run GGUF Models Locally With One Go Binary", "summary": "Janus, an MIT-licensed single Go binary from the Vibra-Ingenn project, launched as a Show HN on October 1 with 42+ points, wrapping llama.cpp's Vulkan backend to run GGUF models on AMD, Intel, or NVIDIA GPUs with CPU fallback and no Python, Docker, or Ollama daemon. Janus exposes an OpenAI-compatible API on localhost via /v1/chat/completions and /v1/models, adds hot-swap model loading through POST /models/load and OpenRouter cloud fallback via INFERENCE_BACKEND=openrouter, and requires Go 1.22+ on Linux while macOS is limited to CPU-only inference. Hacker News commenters noted Intel GPU Vulkan overhead can undercut CPU-only gains depending on driver version, and no published benchmarks compare Janus to vLLM or SGLang.", "body_md": "Setting up local LLM inference has always carried a tax: install Python, configure a virtual environment, wrestle with CUDA drivers, debug ROCm compatibility, then finally load your model. [Janus](https://github.com/Vibra-Ingenn/Janus) eliminates every step between “download binary” and “query your model.” It’s a single MIT-licensed Go binary that wraps llama.cpp’s Vulkan backend, exposes an OpenAI-compatible API on localhost, and runs GGUF models on AMD, Intel, or NVIDIA GPUs with CPU fallback — no Python, no Docker, no Ollama daemon required. The project surfaced as a Show HN on October 1 with 42+ points.\n\n## What Janus Is\n\nJanus is not a service, a registry, or a managed runtime. It’s one executable. The binary ships with pre-compiled llama.cpp Vulkan DLLs on Windows, so there’s no C++ build step. Drop a `.gguf` file in the `models/` directory, point the `JANUS_MODEL_PATH` environment variable at it, and the server is up. The API surface is intentionally minimal: `/v1/chat/completions` and `/v1/models` follow the OpenAI spec exactly, which means any OpenAI-compatible client works without modification.\n\nBeyond standard inference, Janus adds two features worth noting. First, hot-swap: you can change the active model without restarting the server by calling `POST /models/load` with a new path. Second, cloud fallback: set `INFERENCE_BACKEND=openrouter` to route requests to [OpenRouter](https://openrouter.ai) when your local GPU is saturated or you want to switch to a larger model. That’s a genuinely useful escape hatch for development workflows where you prototype locally but occasionally need a heavier model.\n\n## Getting Started\n\nThe build process is minimal. On Windows, `build.ps1` downloads the pre-built llama.cpp Vulkan DLLs and compiles `dist\\janus.exe` in one step. Linux requires Go 1.22+ and a Vulkan-capable driver. macOS users get CPU-only inference — Vulkan support on macOS is not guaranteed and Ollama remains the better choice there.\n\n```\n# Clone and build\ngit clone https://github.com/Vibra-Ingenn/Janus\ncd Janus\n./build.ps1          # Windows\n# make build         # Linux\n\n# Drop your model in models/\ncp ~/Downloads/llama3.2-3b-q4_k_m.gguf models/\n\n# Run the server\nJANUS_MODEL_PATH=models/llama3.2-3b-q4_k_m.gguf ./janus\n\n# Query it — identical to OpenAI API calls\ncurl http://localhost:8080/v1/chat/completions \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\"model\":\"local\",\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}]}'\n```\n\nVRAM tuning uses two variables: `JANUS_GPU_LAYERS` (`-1` offloads all layers to GPU, `0` forces CPU) and `JANUS_VRAM_CEILING_MB` sets a budget cap in MiB. Hit `/engine/status` to see live VRAM usage and the active backend.\n\n## Janus vs Ollama vs llama.cpp Server\n\n[Ollama](https://ollama.com) has become the default local inference choice for most developers, and it deserves that position. Its built-in model registry, `ollama pull` workflow, and Metal acceleration on Apple Silicon are genuinely polished. Janus is not trying to replace that. The differences are narrow but meaningful:\n\n|  | Janus | Ollama | llama.cpp server | \n|---|---|---|---|\n| Install | Single binary | Single binary + daemon | Build from C++ | \n| GPU backends | Vulkan | CUDA / ROCm / Metal | CUDA, ROCm, Metal, Vulkan | \n| Apple Silicon | CPU only | Metal (fast) | Metal (fast) | \n| Hot-swap models | Yes (via API) | No (restart required) | No | \n| Cloud fallback | OpenRouter | None | None | \n| Model registry | No (raw .gguf) | Built-in | No | \n| API | OpenAI-compat | OpenAI-compat | OpenAI-compat | \n\nThe Hacker News thread flagged two real concerns. Intel GPU users report that Vulkan overhead can undercut performance gains versus CPU-only inference, depending on driver version. And there are no published benchmarks comparing Janus to [vLLM](https://docs.vllm.ai) or SGLang — the target audience is single-user local inference, not production serving, but that should be stated explicitly.\n\n## Where Janus Fits\n\nThree scenarios favor Janus over Ollama. One: CI/CD pipelines where dropping a binary and setting two environment variables is cleaner than running an Ollama daemon in a container. Two: Go projects where adding a Python runtime or a separate Ollama service creates real dependency friction. Three: mixed-GPU development teams where one developer has NVIDIA, another has AMD, and the Vulkan backend normalizes the difference without separate driver paths.\n\nFor Apple Silicon users, Ollama’s Metal integration is faster and Janus adds nothing. For high-concurrency production use, neither tool applies — vLLM is the answer. And for developers who want a model registry with one-command downloads, Ollama’s workflow is smoother.\n\nJanus is a precision tool. Its value proposition is exactly what it says: one binary, no runtime dependencies, local GGUF inference, OpenAI-compatible API. For the workflows where that matters, it’s hard to beat.", "url": "https://wpnews.pro/news/janus-run-gguf-models-locally-with-one-go-binary", "canonical_source": "https://byteiota.com/janus-go-gguf-vulkan-llm/", "published_at": "2026-10-02 06:11:58+00:00", "updated_at": "2026-10-02 06:15:18.630298+00:00", "lang": "en", "topics": ["ai-tools", "large-language-models", "ai-infrastructure", "developer-tools"], "entities": ["Janus", "Vibra-Ingenn", "llama.cpp", "OpenRouter", "Ollama", "vLLM", "SGLang", "Hacker News"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/janus-run-gguf-models-locally-with-one-go-binary", "markdown": "https://wpnews.pro/news/janus-run-gguf-models-locally-with-one-go-binary.md", "text": "https://wpnews.pro/news/janus-run-gguf-models-locally-with-one-go-binary.txt", "jsonld": "https://wpnews.pro/news/janus-run-gguf-models-locally-with-one-go-binary.jsonld"}}