{"slug": "running-a-180b-model-on-a-gaming-laptop-pocket-darwin-180b", "title": "Running a 180B Model on a Gaming Laptop: POCKET-Darwin-180B", "summary": "VIDRAFT released POCKET-Darwin-180B, a 180-billion-parameter sparse Mixture-of-Experts model that runs on a consumer gaming laptop with 8GB of VRAM and 32GB of system memory by streaming weights from SSD through the llama.cpp runtime. The company says the goal is not cloud-API speed but offline reach, targeting air-gapped and regulated environments such as defense, healthcare, and legal work where data cannot leave isolated networks.", "body_md": "VIDRAFT released POCKET-Darwin-180B, a 180B-parameter model that runs on a consumer gaming laptop with 8GB of VRAM and 32GB of system memory. It uses a sparse Mixture-of-Experts design, SSD streaming, and the llama.cpp runtime. The point is not to beat a cloud API on speed. The point is to reach places a cloud API cannot go: air-gapped and regulated systems.\n\nA Chinese AI outlet recently covered the model under the headline \"a 180B model moves into a gaming laptop.\" It framed the result as a large drop in the hardware barrier for running big models locally. That framing matches our design goal. We built POCKET-Darwin-180B for offline, isolated machines, not for a data center.\n\nThree ideas work together.\n\nTogether these turn \"needs a server\" into \"runs on the laptop you already own.\"\n\nCloud APIs are easy when you can send data out. Many buyers cannot. Defense, healthcare, and legal work often run on isolated networks by law or policy. For them, a model that runs fully offline on local hardware is not a convenience. It is the only option.\n\nThis is a different market from the cloud API race. The question is not \"who is fastest per token.\" The question is \"who runs at all, inside the wall, on hardware the customer already has.\"\n\nLocal and offline is a real deployment target, not a demo. A 180B model on a gaming laptop shows that the frontier is not only in the cloud. It is also on the disconnected machine in a secure room.\n\nPer token, a sparse MoE activates a small subset of experts. The full 180B parameters exist, but each token uses a fraction of them.\n\nYes, local inference on a laptop is slower than a tuned cloud endpoint. The trade is latency for reach: it runs where the cloud cannot.\n\nThe reported setup is a gaming laptop with 8GB of VRAM and 32GB of system memory, with weights streamed from SSD.\n\nAir-gapped and regulated environments: defense, healthcare, legal, and any site that cannot send data to an external API.", "url": "https://wpnews.pro/news/running-a-180b-model-on-a-gaming-laptop-pocket-darwin-180b", "canonical_source": "https://dev.to/ai_openfree_b23025ef075cf/running-a-180b-model-on-a-gaming-laptop-pocket-darwin-180b-1m0l", "published_at": "2026-10-08 01:59:27+00:00", "updated_at": "2026-10-08 02:17:54.817766+00:00", "lang": "en", "topics": ["large-language-models", "ai-infrastructure", "ai-tools", "mlops"], "entities": ["VIDRAFT", "POCKET-Darwin-180B", "llama.cpp"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/running-a-180b-model-on-a-gaming-laptop-pocket-darwin-180b", "markdown": "https://wpnews.pro/news/running-a-180b-model-on-a-gaming-laptop-pocket-darwin-180b.md", "text": "https://wpnews.pro/news/running-a-180b-model-on-a-gaming-laptop-pocket-darwin-180b.txt", "jsonld": "https://wpnews.pro/news/running-a-180b-model-on-a-gaming-laptop-pocket-darwin-180b.jsonld"}}