TL;DR: VIDRAFT has released POCKET-Darwin-180B-GGUF, a 4-bit quantized, GGUF-format version of their 180B Mixture-of-Experts model that runs on consumer hardware — including a gaming laptop with 8 GB VRAM and 32 GB RAM. By combining MoE sparse activation with llama.cpp-based SSD streaming, only ~3B parameters are computed per token at inference time, slashing hardware requirements by roughly 250× compared to the full-precision server setup. Developers working on air-gapped or on-device deployments should pay attention.
POCKET-Darwin-180B-GGUF is a compressed, consumer-runnable edition of VIDRAFT's Darwin-180B-RSI foundation model — a 180-billion-parameter model built on top of their Qwen3.8-Flash-Next base. The original Darwin-180B-RSI weighs in at 360 GB at BF16 precision and conventionally requires a multi-GPU server to serve.
The POCKET variant addresses that head-on:
The model name follows VIDRAFT's stated product strategy: RSI (Recursive Self-Improvement) for capability, POCKET for deployability.
The key insight is that Darwin-180B was architected as a Mixture-of-Experts (MoE) model from the ground up, which makes aggressive compression tractable without destroying capability.
Here's the conceptual stack:
Sparse MoE architecture: The full model contains 512 expert sub-networks. At each token, only 10 experts are activated, meaning the effective compute per token corresponds to roughly ~3B active parameters, not the full 180B. Total parameter count and active parameter count are very different numbers in MoE — this is the core of why quantization here is less lossy than it would be on a dense model.
4-bit quantization: Model weights are converted from BF16 to 4-bit integer representations. This shrinks the storage footprint from 360 GB to 111 GB. The MoE routing structure is preserved — the model still selects 10 out of 512 experts per token, just using quantized weights.
GGUF format + llama.cpp SSD streaming: Rather than the entire 111 GB into RAM at inference time, the system uses llama.cpp's ability to stream weights from an SSD on demand. Only the expert weights needed for the current token need to be resident in memory at any given moment. This is what makes the 8 GB VRAM + 32 GB RAM laptop configuration viable — the SSD becomes an extension of the memory hierarchy.
The combination of these three techniques — sparse activation (MoE), aggressive quantization, and demand-paged weight streaming — is what lets a model with 180B total parameters run in environments where a 7B dense model would typically max out.
VIDRAFT published the following figures from their own testing (self-reported):
| Hardware config | Inference speed | Notes |
|---|---|---|
| Server CPU only, 1× CPU (16 threads) | 18.4 – 21.0 tokens/sec | Memory usage: 78.8 GB |
| Mini-PC with 128 GB RAM | Full model in RAM | GPU-free, CPU-only inference |
| Gaming laptop (8 GB VRAM + 32 GB RAM) | 4.17 tokens/sec | SSD streaming mode |
Accuracy retention after quantization (MMLU-Pro, 2,000 questions, matched items):
No measurable accuracy degradation was observed on MMLU-Pro under their evaluation methodology. This is a self-reported result on a matched question set — independent replication is always encouraged.
The model is publicly available on Hugging Face and ModelScope in GGUF format. You need a recent build of llama.cpp and approximately 111 GB of free storage.
Download via Hugging Face CLI:
pip install huggingface_hub
huggingface-cli download VIDRAFT/POCKET-Darwin-180B-GGUF --local-dir ./pocket-darwin-180b
Once downloaded, run with llama.cpp (generic invocation — check the model card for the exact recommended flags):
./llama-cli -m ./pocket-darwin-180b/<model-file>.gguf \
--n-gpu-layers <N> \
-p "Your prompt here"
Set --n-gpu-layers 0 for pure CPU inference, or a positive integer to offload layers to your GPU. For SSD-streaming on low-RAM hardware, refer to the llama.cpp documentation on mmap and context management.
Check the model card on Hugging Face for the exact file names, recommended context length, and chat template.
Q: Why does a 180B model only need ~3B parameters' worth of compute per token?
A: This is the defining property of Mixture-of-Experts. The 180B count is the total parameter budget across all 512 expert networks. At inference, a learned router selects only 10 of those 512 experts per token — so the actual floating-point operations per token correspond to roughly 3B active parameters. You pay the storage cost of 180B but the compute cost of ~3B.
Q: Does 4-bit quantization meaningfully hurt quality for a model this size?
A: Based on VIDRAFT's published MMLU-Pro evaluation, accuracy was identical before and after quantization on their 2,000-item matched test set. That said, quantization effects can vary by task type and prompt style — benchmarking on your specific workload is always advisable.
Q: Is this useful for enterprise air-gapped deployments?
A: Yes, that appears to be a primary intended use case. Because the model runs entirely locally via llama.cpp with no cloud dependency, it suits environments — enterprise, government, regulated industries — where data cannot leave the local network.
Q: What is the RSI component referenced in the base model name?
A: RSI stands for Recursive Self-Improvement, VIDRAFT's stated training methodology for improving model capability. The POCKET release is the inference/deployment layer on top of that capability; the RSI process relates to how Darwin-180B was trained, not how POCKET-Darwin-180B-GGUF is quantized or served.
Originally reported by AI타임스 (2026-10-06) — source article.