cd /news/ai-tools/janus-run-gguf-models-locally-with-o… · home › topics › ai-tools › article
[ARTICLE · art-143696] src=byteiota.com ↗ pub= topic=ai-tools verified=true sentiment=↑ positive

Janus: Run GGUF Models Locally With One Go Binary

Janus, an MIT-licensed single Go binary from the Vibra-Ingenn project, launched as a Show HN on October 1 with 42+ points, wrapping llama.cpp's Vulkan backend to run GGUF models on AMD, Intel, or NVIDIA GPUs with CPU fallback and no Python, Docker, or Ollama daemon. Janus exposes an OpenAI-compatible API on localhost via /v1/chat/completions and /v1/models, adds hot-swap model loading through POST /models/load and OpenRouter cloud fallback via INFERENCE_BACKEND=openrouter, and requires Go 1.22+ on Linux while macOS is limited to CPU-only inference. Hacker News commenters noted Intel GPU Vulkan overhead can undercut CPU-only gains depending on driver version, and no published benchmarks compare Janus to vLLM or SGLang.

read4 min views3 publishedOct 2, 2026
Janus: Run GGUF Models Locally With One Go Binary
Image: Byteiota (auto-discovered)

Setting up local LLM inference has always carried a tax: install Python, configure a virtual environment, wrestle with CUDA drivers, debug ROCm compatibility, then finally load your model. Janus eliminates every step between “download binary” and “query your model.” It’s a single MIT-licensed Go binary that wraps llama.cpp’s Vulkan backend, exposes an OpenAI-compatible API on localhost, and runs GGUF models on AMD, Intel, or NVIDIA GPUs with CPU fallback — no Python, no Docker, no Ollama daemon required. The project surfaced as a Show HN on October 1 with 42+ points.

What Janus Is #

Janus is not a service, a registry, or a managed runtime. It’s one executable. The binary ships with pre-compiled llama.cpp Vulkan DLLs on Windows, so there’s no C++ build step. Drop a .gguf file in the models/ directory, point the JANUS_MODEL_PATH environment variable at it, and the server is up. The API surface is intentionally minimal: /v1/chat/completions and /v1/models follow the OpenAI spec exactly, which means any OpenAI-compatible client works without modification.

Beyond standard inference, Janus adds two features worth noting. First, hot-swap: you can change the active model without restarting the server by calling POST /models/load with a new path. Second, cloud fallback: set INFERENCE_BACKEND=openrouter to route requests to OpenRouter when your local GPU is saturated or you want to switch to a larger model. That’s a genuinely useful escape hatch for development workflows where you prototype locally but occasionally need a heavier model.

Getting Started #

The build process is minimal. On Windows, build.ps1 downloads the pre-built llama.cpp Vulkan DLLs and compiles dist\janus.exe in one step. Linux requires Go 1.22+ and a Vulkan-capable driver. macOS users get CPU-only inference — Vulkan support on macOS is not guaranteed and Ollama remains the better choice there.

git clone https://github.com/Vibra-Ingenn/Janus
cd Janus
./build.ps1          # Windows

cp ~/Downloads/llama3.2-3b-q4_k_m.gguf models/

JANUS_MODEL_PATH=models/llama3.2-3b-q4_k_m.gguf ./janus

curl http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"local","messages":[{"role":"user","content":"Hello"}]}'

VRAM tuning uses two variables: JANUS_GPU_LAYERS (-1 offloads all layers to GPU, 0 forces CPU) and JANUS_VRAM_CEILING_MB sets a budget cap in MiB. Hit /engine/status to see live VRAM usage and the active backend.

Janus vs Ollama vs llama.cpp Server #

Ollama has become the default local inference choice for most developers, and it deserves that position. Its built-in model registry, ollama pull workflow, and Metal acceleration on Apple Silicon are genuinely polished. Janus is not trying to replace that. The differences are narrow but meaningful:

Janus Ollama llama.cpp server
Install Single binary Single binary + daemon Build from C++
GPU backends Vulkan CUDA / ROCm / Metal CUDA, ROCm, Metal, Vulkan
Apple Silicon CPU only Metal (fast) Metal (fast)
Hot-swap models Yes (via API) No (restart required) No
Cloud fallback OpenRouter None None
Model registry No (raw .gguf) Built-in No
API OpenAI-compat OpenAI-compat OpenAI-compat

The Hacker News thread flagged two real concerns. Intel GPU users report that Vulkan overhead can undercut performance gains versus CPU-only inference, depending on driver version. And there are no published benchmarks comparing Janus to vLLM or SGLang — the target audience is single-user local inference, not production serving, but that should be stated explicitly.

Where Janus Fits #

Three scenarios favor Janus over Ollama. One: CI/CD pipelines where dropping a binary and setting two environment variables is cleaner than running an Ollama daemon in a container. Two: Go projects where adding a Python runtime or a separate Ollama service creates real dependency friction. Three: mixed-GPU development teams where one developer has NVIDIA, another has AMD, and the Vulkan backend normalizes the difference without separate driver paths.

For Apple Silicon users, Ollama’s Metal integration is faster and Janus adds nothing. For high-concurrency production use, neither tool applies — vLLM is the answer. And for developers who want a model registry with one-command downloads, Ollama’s workflow is smoother.

Janus is a precision tool. Its value proposition is exactly what it says: one binary, no runtime dependencies, local GGUF inference, OpenAI-compatible API. For the workflows where that matters, it’s hard to beat.

── more in #ai-tools 4 stories · sorted by recency
── more on @janus 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/janus-run-gguf-model…] indexed:0 read:4min 2026-10-02 · —