Leviathan — a local, persistent AI runtime for Windows + NVIDIA (no server, no cloud, no account). Feedback welcome Solo developer OmegaVR (GitHub handle CuppaTea1983) released Leviathan, a local Windows AI runtime that loads GGUF models directly into NVIDIA GPU VRAM in-process with no server, Docker, Python, CUDA toolkit, or API key required. Leviathan uses hand-written CUDA kernels validated byte-exact against llama.cpp on Llama/Mistral, Qwen2.5, and DeepSeek-V2 (MLA + MoE), converting each .gguf once to a native .lev file that is mmapped to VRAM, and adds persistent .fqm memory, ANI knowledge routing, a layered security membrane, and an OpenAI-compatible local bridge. The developer states Leviathan is NVIDIA-only (GTX 16 / RTX 20 or newer), is not faster than llama.cpp on raw throughput, and is set to be released soon under CC BY-NC 4.0, with a Vulkan backend for AMD/Intel planned but not shipped. Hi all — I’ve been building this solo for a while and it’s finally at a state worth sharing. It’s part of a project I call Sovereign; Leviathan is the public Windows app. What it is: you open the app, pick a GGUF convert it, and then talk to it — the model loads straight into your GPU’s VRAM and runs in-process. No local server, no Docker, no Python or CUDA toolkit to install, no API key. Just a recent NVIDIA driver. It’s its own CUDA inference engine hand-written kernels , and I’ve validated it byte-exact against llama.cpp on the models it runs — Llama/Mistral, Qwen2.5, and DeepSeek-V2 MLA + MoE . It converts a .gguf to a native .lev once, then mmaps it to VRAM. What makes it different is everything around the model the host/guest idea — Leviathan is the host, the model is the guest : Persistent memory .fqm — the model remembers across sessions; close the app and reopen, it’s still there. Per-tab scoped so a coding workspace and a chat don’t bleed into each other. Knowledge routing ANI — teach it things without retraining. Knowledge is kept as text on a shared embedding space and retrieved on relevance — drained from a model’s own confident answers, folded from your chat logs / documents / audio / screen / voice. I tested weight-level transfer thoroughly and it doesn’t survive across architectures — text + retrieval is the honest carrier, so that’s what it uses. A security membrane — a layered input guard prompt-injection, malicious-code, an always-on floor, a model-file integrity scan before a model reaches VRAM, and knowledge-bank poisoning checks . The sovereignty cuts both ways: nothing leaves your machine, and nothing hostile gets in. A local bridge — your own programs a game engine, a script can use the model Leviathan is hosting, over an OpenAI-compatible endpoint. Honest boundaries, because I’d rather you trust the page than catch me overclaiming: NVIDIA-only right now GTX 16 / RTX 20 or newer . A Vulkan backend for AMD/Intel is planned, not shipped. It is not faster than llama.cpp on raw throughput — llama.cpp is the most optimised engine out there. Leviathan’s edge is architectural in-process, no middleman, + the memory/knowledge/security layers , not tokens/sec. Web-free by default; if you connect a hosted model e.g. Grok that’s clearly separate and obviously leaves your machine — everything on your own GPU stays local. Leviathan is set to be released soon, just ironing out UI mostly. Space Includes: How it works full guide Leviathan - a Hugging Face Space by Omega-Dev https://huggingface.co/spaces/Omega-Dev/Leviathan GitHub: CuppaTea1983/Sovereign: Sovereign — A Geometric Substrate for Persistent Intelligence A unified cognitive substrate architecture integrating memory, compression, geometry, and identity. https://github.com/CuppaTea1983/Sovereign Non-commercial CC BY-NC 4.0 . I’d genuinely love eyes on the knowledge-routing / persistence design and the security layering — tell me where it’s weak. — OmegaVR / @CuppaTea1983