Hi all — I’ve been building this solo for a while and it’s finally at a state worth sharing. It’s part of a project I call Sovereign; Leviathan is the public Windows app.
What it is: you open the app, pick a GGUF convert it, and then talk to it — the model loads straight into your GPU’s VRAM and runs in-process. No local server, no Docker, no Python or CUDA toolkit to install, no API key. Just a recent NVIDIA driver. It’s its own CUDA inference engine (hand-written kernels), and I’ve validated it byte-exact against llama.cpp on the models it runs — Llama/Mistral, Qwen2.5, and DeepSeek-V2 (MLA + MoE). It converts a .gguf to a native .lev once, then mmaps it to VRAM.
What makes it different is everything around the model (the host/guest idea — Leviathan is the host, the model is the guest):
Persistent memory (.fqm) — the model remembers across sessions; close the app and reopen, it’s still there. Per-tab scoped so a coding workspace and a chat don’t bleed into each other.
Knowledge routing (ANI) — teach it things without retraining. Knowledge is kept as text on a shared embedding space and retrieved on relevance — drained from a model’s own confident answers, folded from your chat logs / documents / audio / screen / voice. (I tested weight-level transfer thoroughly and it doesn’t survive across architectures — text + retrieval is the honest carrier, so that’s what it uses.)
A security membrane — a layered input guard (prompt-injection, malicious-code, an always-on floor, a model-file integrity scan before a model reaches VRAM, and knowledge-bank poisoning checks). The sovereignty cuts both ways: nothing leaves your machine, and nothing hostile gets in.
A local bridge — your own programs (a game engine, a script) can use the model Leviathan is hosting, over an OpenAI-compatible endpoint.
Honest boundaries, because I’d rather you trust the page than catch me overclaiming:
NVIDIA-only right now (GTX 16 / RTX 20 or newer). A Vulkan backend for AMD/Intel is planned, not shipped.
It is not faster than llama.cpp on raw throughput — llama.cpp is the most optimised engine out there. Leviathan’s edge is architectural (in-process, no middleman, + the memory/knowledge/security layers), not tokens/sec.
Web-free by default; if you connect a hosted model (e.g. Grok) that’s clearly separate and obviously leaves your machine — everything on your own GPU stays local.
Leviathan is set to be released soon, just ironing out UI mostly.
Space Includes: How it works (full guide) [Leviathan - a Hugging Face Space by Omega-Dev](https://huggingface.co/spaces/Omega-Dev/Leviathan)
GitHub: [CuppaTea1983/Sovereign: Sovereign — A Geometric Substrate for Persistent Intelligence A unified cognitive substrate architecture integrating memory, compression, geometry, and identity.](https://github.com/CuppaTea1983/Sovereign)
Non-commercial (CC BY-NC 4.0). I’d genuinely love eyes on the knowledge-routing / persistence design and the security layering — tell me where it’s weak. — OmegaVR / @CuppaTea1983