cd /news/large-language-models/running-a-180b-parameter-moe-model-o… · home › topics › large-language-models › article
[ARTICLE · art-146418] src=dev.to ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

Running a 180B-Parameter MoE Model on a Gaming Laptop: VIDRAFT's POCKET-Darwin-180B-GGUF

VIDRAFT released POCKET-Darwin-180B-GGUF, a 4-bit quantized GGUF build of its 180B-parameter Mixture-of-Experts Darwin-180B-RSI model that runs on consumer hardware, including a gaming laptop with 8 GB VRAM and 32 GB RAM. By activating only 10 of 512 experts per token (~3B active parameters) and streaming weights from SSD via llama.cpp, the release cuts storage from 360 GB at BF16 to 111 GB and reports 4.17 tokens/sec on the laptop configuration with no measurable MMLU-Pro degradation in the company's own testing.

by read4 min views1 publishedOct 6, 2026

TL;DR: VIDRAFT has released POCKET-Darwin-180B-GGUF, a 4-bit quantized, GGUF-format version of their 180B Mixture-of-Experts model that runs on consumer hardware — including a gaming laptop with 8 GB VRAM and 32 GB RAM. By combining MoE sparse activation with llama.cpp-based SSD streaming, only ~3B parameters are computed per token at inference time, slashing hardware requirements by roughly 250× compared to the full-precision server setup. Developers working on air-gapped or on-device deployments should pay attention.

POCKET-Darwin-180B-GGUF is a compressed, consumer-runnable edition of VIDRAFT's Darwin-180B-RSI foundation model — a 180-billion-parameter model built on top of their Qwen3.8-Flash-Next base. The original Darwin-180B-RSI weighs in at 360 GB at BF16 precision and conventionally requires a multi-GPU server to serve.

The POCKET variant addresses that head-on:

The model name follows VIDRAFT's stated product strategy: RSI (Recursive Self-Improvement) for capability, POCKET for deployability.

The key insight is that Darwin-180B was architected as a Mixture-of-Experts (MoE) model from the ground up, which makes aggressive compression tractable without destroying capability.

Here's the conceptual stack:

Sparse MoE architecture: The full model contains 512 expert sub-networks. At each token, only 10 experts are activated, meaning the effective compute per token corresponds to roughly ~3B active parameters, not the full 180B. Total parameter count and active parameter count are very different numbers in MoE — this is the core of why quantization here is less lossy than it would be on a dense model.

4-bit quantization: Model weights are converted from BF16 to 4-bit integer representations. This shrinks the storage footprint from 360 GB to 111 GB. The MoE routing structure is preserved — the model still selects 10 out of 512 experts per token, just using quantized weights.

GGUF format + llama.cpp SSD streaming: Rather than the entire 111 GB into RAM at inference time, the system uses llama.cpp's ability to stream weights from an SSD on demand. Only the expert weights needed for the current token need to be resident in memory at any given moment. This is what makes the 8 GB VRAM + 32 GB RAM laptop configuration viable — the SSD becomes an extension of the memory hierarchy.

The combination of these three techniques — sparse activation (MoE), aggressive quantization, and demand-paged weight streaming — is what lets a model with 180B total parameters run in environments where a 7B dense model would typically max out.

VIDRAFT published the following figures from their own testing (self-reported):

Hardware config Inference speed Notes
Server CPU only, 1× CPU (16 threads) 18.4 – 21.0 tokens/sec Memory usage: 78.8 GB
Mini-PC with 128 GB RAM Full model in RAM GPU-free, CPU-only inference
Gaming laptop (8 GB VRAM + 32 GB RAM) 4.17 tokens/sec SSD streaming mode

Accuracy retention after quantization (MMLU-Pro, 2,000 questions, matched items):

No measurable accuracy degradation was observed on MMLU-Pro under their evaluation methodology. This is a self-reported result on a matched question set — independent replication is always encouraged.

The model is publicly available on Hugging Face and ModelScope in GGUF format. You need a recent build of llama.cpp and approximately 111 GB of free storage.

Download via Hugging Face CLI:

pip install huggingface_hub
huggingface-cli download VIDRAFT/POCKET-Darwin-180B-GGUF --local-dir ./pocket-darwin-180b

Once downloaded, run with llama.cpp (generic invocation — check the model card for the exact recommended flags):

./llama-cli -m ./pocket-darwin-180b/<model-file>.gguf \
  --n-gpu-layers <N> \
  -p "Your prompt here"

Set --n-gpu-layers 0 for pure CPU inference, or a positive integer to offload layers to your GPU. For SSD-streaming on low-RAM hardware, refer to the llama.cpp documentation on mmap and context management.

Check the model card on Hugging Face for the exact file names, recommended context length, and chat template.

Q: Why does a 180B model only need ~3B parameters' worth of compute per token?

A: This is the defining property of Mixture-of-Experts. The 180B count is the total parameter budget across all 512 expert networks. At inference, a learned router selects only 10 of those 512 experts per token — so the actual floating-point operations per token correspond to roughly 3B active parameters. You pay the storage cost of 180B but the compute cost of ~3B.

Q: Does 4-bit quantization meaningfully hurt quality for a model this size?

A: Based on VIDRAFT's published MMLU-Pro evaluation, accuracy was identical before and after quantization on their 2,000-item matched test set. That said, quantization effects can vary by task type and prompt style — benchmarking on your specific workload is always advisable.

Q: Is this useful for enterprise air-gapped deployments?

A: Yes, that appears to be a primary intended use case. Because the model runs entirely locally via llama.cpp with no cloud dependency, it suits environments — enterprise, government, regulated industries — where data cannot leave the local network.

Q: What is the RSI component referenced in the base model name?

A: RSI stands for Recursive Self-Improvement, VIDRAFT's stated training methodology for improving model capability. The POCKET release is the inference/deployment layer on top of that capability; the RSI process relates to how Darwin-180B was trained, not how POCKET-Darwin-180B-GGUF is quantized or served.

Originally reported by AI타임스 (2026-10-06) — source article.

── more in #large-language-models 4 stories · sorted by recency
── more on @vidraft 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/running-a-180b-param…] indexed:0 read:4min 2026-10-06 · —