# Janus: Run GGUF Models Locally With One Go Binary

> Source: <https://byteiota.com/janus-go-gguf-vulkan-llm/>
> Published: 2026-10-02 06:11:58+00:00

Setting up local LLM inference has always carried a tax: install Python, configure a virtual environment, wrestle with CUDA drivers, debug ROCm compatibility, then finally load your model. [Janus](https://github.com/Vibra-Ingenn/Janus) eliminates every step between “download binary” and “query your model.” It’s a single MIT-licensed Go binary that wraps llama.cpp’s Vulkan backend, exposes an OpenAI-compatible API on localhost, and runs GGUF models on AMD, Intel, or NVIDIA GPUs with CPU fallback — no Python, no Docker, no Ollama daemon required. The project surfaced as a Show HN on October 1 with 42+ points.

## What Janus Is

Janus is not a service, a registry, or a managed runtime. It’s one executable. The binary ships with pre-compiled llama.cpp Vulkan DLLs on Windows, so there’s no C++ build step. Drop a `.gguf` file in the `models/` directory, point the `JANUS_MODEL_PATH` environment variable at it, and the server is up. The API surface is intentionally minimal: `/v1/chat/completions` and `/v1/models` follow the OpenAI spec exactly, which means any OpenAI-compatible client works without modification.

Beyond standard inference, Janus adds two features worth noting. First, hot-swap: you can change the active model without restarting the server by calling `POST /models/load` with a new path. Second, cloud fallback: set `INFERENCE_BACKEND=openrouter` to route requests to [OpenRouter](https://openrouter.ai) when your local GPU is saturated or you want to switch to a larger model. That’s a genuinely useful escape hatch for development workflows where you prototype locally but occasionally need a heavier model.

## Getting Started

The build process is minimal. On Windows, `build.ps1` downloads the pre-built llama.cpp Vulkan DLLs and compiles `dist\janus.exe` in one step. Linux requires Go 1.22+ and a Vulkan-capable driver. macOS users get CPU-only inference — Vulkan support on macOS is not guaranteed and Ollama remains the better choice there.

```
# Clone and build
git clone https://github.com/Vibra-Ingenn/Janus
cd Janus
./build.ps1          # Windows
# make build         # Linux

# Drop your model in models/
cp ~/Downloads/llama3.2-3b-q4_k_m.gguf models/

# Run the server
JANUS_MODEL_PATH=models/llama3.2-3b-q4_k_m.gguf ./janus

# Query it — identical to OpenAI API calls
curl http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"local","messages":[{"role":"user","content":"Hello"}]}'
```

VRAM tuning uses two variables: `JANUS_GPU_LAYERS` (`-1` offloads all layers to GPU, `0` forces CPU) and `JANUS_VRAM_CEILING_MB` sets a budget cap in MiB. Hit `/engine/status` to see live VRAM usage and the active backend.

## Janus vs Ollama vs llama.cpp Server

[Ollama](https://ollama.com) has become the default local inference choice for most developers, and it deserves that position. Its built-in model registry, `ollama pull` workflow, and Metal acceleration on Apple Silicon are genuinely polished. Janus is not trying to replace that. The differences are narrow but meaningful:

|  | Janus | Ollama | llama.cpp server | 
|---|---|---|---|
| Install | Single binary | Single binary + daemon | Build from C++ | 
| GPU backends | Vulkan | CUDA / ROCm / Metal | CUDA, ROCm, Metal, Vulkan | 
| Apple Silicon | CPU only | Metal (fast) | Metal (fast) | 
| Hot-swap models | Yes (via API) | No (restart required) | No | 
| Cloud fallback | OpenRouter | None | None | 
| Model registry | No (raw .gguf) | Built-in | No | 
| API | OpenAI-compat | OpenAI-compat | OpenAI-compat | 

The Hacker News thread flagged two real concerns. Intel GPU users report that Vulkan overhead can undercut performance gains versus CPU-only inference, depending on driver version. And there are no published benchmarks comparing Janus to [vLLM](https://docs.vllm.ai) or SGLang — the target audience is single-user local inference, not production serving, but that should be stated explicitly.

## Where Janus Fits

Three scenarios favor Janus over Ollama. One: CI/CD pipelines where dropping a binary and setting two environment variables is cleaner than running an Ollama daemon in a container. Two: Go projects where adding a Python runtime or a separate Ollama service creates real dependency friction. Three: mixed-GPU development teams where one developer has NVIDIA, another has AMD, and the Vulkan backend normalizes the difference without separate driver paths.

For Apple Silicon users, Ollama’s Metal integration is faster and Janus adds nothing. For high-concurrency production use, neither tool applies — vLLM is the answer. And for developers who want a model registry with one-command downloads, Ollama’s workflow is smoother.

Janus is a precision tool. Its value proposition is exactly what it says: one binary, no runtime dependencies, local GGUF inference, OpenAI-compatible API. For the workflows where that matters, it’s hard to beat.
