Running a 180B-Parameter MoE Model on a Gaming Laptop: VIDRAFT's POCKET-Darwin-180B-GGUF VIDRAFT released POCKET-Darwin-180B-GGUF, a 4-bit quantized GGUF build of its 180B-parameter Mixture-of-Experts Darwin-180B-RSI model that runs on consumer hardware, including a gaming laptop with 8 GB VRAM and 32 GB RAM. By activating only 10 of 512 experts per token (~3B active parameters) and streaming weights from SSD via llama.cpp, the release cuts storage from 360 GB at BF16 to 111 GB and reports 4.17 tokens/sec on the laptop configuration with no measurable MMLU-Pro degradation in the company's own testing. TL;DR: VIDRAFT has released POCKET-Darwin-180B-GGUF, a 4-bit quantized, GGUF-format version of their 180B Mixture-of-Experts model that runs on consumer hardware — including a gaming laptop with 8 GB VRAM and 32 GB RAM. By combining MoE sparse activation with llama.cpp-based SSD streaming, only ~3B parameters are computed per token at inference time, slashing hardware requirements by roughly 250× compared to the full-precision server setup. Developers working on air-gapped or on-device deployments should pay attention. POCKET-Darwin-180B-GGUF is a compressed, consumer-runnable edition of VIDRAFT's Darwin-180B-RSI foundation model — a 180-billion-parameter model built on top of their Qwen3.8-Flash-Next base. The original Darwin-180B-RSI weighs in at 360 GB at BF16 precision and conventionally requires a multi-GPU server to serve. The POCKET variant addresses that head-on: The model name follows VIDRAFT's stated product strategy: RSI Recursive Self-Improvement for capability, POCKET for deployability. The key insight is that Darwin-180B was architected as a Mixture-of-Experts MoE model from the ground up, which makes aggressive compression tractable without destroying capability. Here's the conceptual stack: Sparse MoE architecture: The full model contains 512 expert sub-networks. At each token, only 10 experts are activated , meaning the effective compute per token corresponds to roughly ~3B active parameters , not the full 180B. Total parameter count and active parameter count are very different numbers in MoE — this is the core of why quantization here is less lossy than it would be on a dense model. 4-bit quantization: Model weights are converted from BF16 to 4-bit integer representations. This shrinks the storage footprint from 360 GB to 111 GB. The MoE routing structure is preserved — the model still selects 10 out of 512 experts per token, just using quantized weights. GGUF format + llama.cpp SSD streaming: Rather than loading the entire 111 GB into RAM at inference time, the system uses llama.cpp's ability to stream weights from an SSD on demand. Only the expert weights needed for the current token need to be resident in memory at any given moment. This is what makes the 8 GB VRAM + 32 GB RAM laptop configuration viable — the SSD becomes an extension of the memory hierarchy. The combination of these three techniques — sparse activation MoE , aggressive quantization, and demand-paged weight streaming — is what lets a model with 180B total parameters run in environments where a 7B dense model would typically max out. VIDRAFT published the following figures from their own testing self-reported : | Hardware config | Inference speed | Notes | |---|---|---| | Server CPU only, 1× CPU 16 threads | 18.4 – 21.0 tokens/sec | Memory usage: 78.8 GB | | Mini-PC with 128 GB RAM | Full model in RAM | GPU-free, CPU-only inference | | Gaming laptop 8 GB VRAM + 32 GB RAM | 4.17 tokens/sec | SSD streaming mode | Accuracy retention after quantization MMLU-Pro, 2,000 questions, matched items : No measurable accuracy degradation was observed on MMLU-Pro under their evaluation methodology. This is a self-reported result on a matched question set — independent replication is always encouraged. The model is publicly available on Hugging Face and ModelScope in GGUF format. You need a recent build of llama.cpp and approximately 111 GB of free storage. Download via Hugging Face CLI: pip install huggingface hub huggingface-cli download VIDRAFT/POCKET-Darwin-180B-GGUF --local-dir ./pocket-darwin-180b Once downloaded, run with llama.cpp generic invocation — check the model card for the exact recommended flags : ./llama-cli -m ./pocket-darwin-180b/