Setting up local LLM inference has always carried a tax: install Python, configure a virtual environment, wrestle with CUDA drivers, debug ROCm compatibility, then finally load your model. Janus eliminates every step between “download binary” and “query your model.” It’s a single MIT-licensed Go binary that wraps llama.cpp’s Vulkan backend, exposes an OpenAI-compatible API on localhost, and runs GGUF models on AMD, Intel, or NVIDIA GPUs with CPU fallback — no Python, no Docker, no Ollama daemon required. The project surfaced as a Show HN on October 1 with 42+ points.
What Janus Is #
Janus is not a service, a registry, or a managed runtime. It’s one executable. The binary ships with pre-compiled llama.cpp Vulkan DLLs on Windows, so there’s no C++ build step. Drop a .gguf file in the models/ directory, point the JANUS_MODEL_PATH environment variable at it, and the server is up. The API surface is intentionally minimal: /v1/chat/completions and /v1/models follow the OpenAI spec exactly, which means any OpenAI-compatible client works without modification.
Beyond standard inference, Janus adds two features worth noting. First, hot-swap: you can change the active model without restarting the server by calling POST /models/load with a new path. Second, cloud fallback: set INFERENCE_BACKEND=openrouter to route requests to OpenRouter when your local GPU is saturated or you want to switch to a larger model. That’s a genuinely useful escape hatch for development workflows where you prototype locally but occasionally need a heavier model.
Getting Started #
The build process is minimal. On Windows, build.ps1 downloads the pre-built llama.cpp Vulkan DLLs and compiles dist\janus.exe in one step. Linux requires Go 1.22+ and a Vulkan-capable driver. macOS users get CPU-only inference — Vulkan support on macOS is not guaranteed and Ollama remains the better choice there.
git clone https://github.com/Vibra-Ingenn/Janus
cd Janus
./build.ps1 # Windows
cp ~/Downloads/llama3.2-3b-q4_k_m.gguf models/
JANUS_MODEL_PATH=models/llama3.2-3b-q4_k_m.gguf ./janus
curl http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"local","messages":[{"role":"user","content":"Hello"}]}'
VRAM tuning uses two variables: JANUS_GPU_LAYERS (-1 offloads all layers to GPU, 0 forces CPU) and JANUS_VRAM_CEILING_MB sets a budget cap in MiB. Hit /engine/status to see live VRAM usage and the active backend.
Janus vs Ollama vs llama.cpp Server #
Ollama has become the default local inference choice for most developers, and it deserves that position. Its built-in model registry, ollama pull workflow, and Metal acceleration on Apple Silicon are genuinely polished. Janus is not trying to replace that. The differences are narrow but meaningful:
| Janus | Ollama | llama.cpp server | |
|---|---|---|---|
| Install | Single binary | Single binary + daemon | Build from C++ |
| GPU backends | Vulkan | CUDA / ROCm / Metal | CUDA, ROCm, Metal, Vulkan |
| Apple Silicon | CPU only | Metal (fast) | Metal (fast) |
| Hot-swap models | Yes (via API) | No (restart required) | No |
| Cloud fallback | OpenRouter | None | None |
| Model registry | No (raw .gguf) | Built-in | No |
| API | OpenAI-compat | OpenAI-compat | OpenAI-compat |
The Hacker News thread flagged two real concerns. Intel GPU users report that Vulkan overhead can undercut performance gains versus CPU-only inference, depending on driver version. And there are no published benchmarks comparing Janus to vLLM or SGLang — the target audience is single-user local inference, not production serving, but that should be stated explicitly.
Where Janus Fits #
Three scenarios favor Janus over Ollama. One: CI/CD pipelines where dropping a binary and setting two environment variables is cleaner than running an Ollama daemon in a container. Two: Go projects where adding a Python runtime or a separate Ollama service creates real dependency friction. Three: mixed-GPU development teams where one developer has NVIDIA, another has AMD, and the Vulkan backend normalizes the difference without separate driver paths.
For Apple Silicon users, Ollama’s Metal integration is faster and Janus adds nothing. For high-concurrency production use, neither tool applies — vLLM is the answer. And for developers who want a model registry with one-command downloads, Ollama’s workflow is smoother.
Janus is a precision tool. Its value proposition is exactly what it says: one binary, no runtime dependencies, local GGUF inference, OpenAI-compatible API. For the workflows where that matters, it’s hard to beat.