Inside Janus: How Go FFI and Vulkan Bypass Local LLM Infrastructure Stack Janus embeds llama.cpp directly inside a lightweight Go binary with a Vulkan acceleration bridge, eliminating the Python runtime and Docker virtualization layers used by typical local LLM stacks. The single-file gateway combines hot-swappable GGUF model execution with native reasoning tag extraction and targets AMD, Intel, and NVIDIA hardware. Janus strips away Python runtime bloat and Docker virtualization, embedding llama.cpp directly inside a lightweight Go binary with a Vulkan acceleration bridge. By uniting hot-swappable GGUF model execution with native reasoning tag extraction, it offers a single-file gateway for local inference across AMD, Intel, and NVIDIA hardware.