Kimi K3 Open Weights Drop Sunday — 2.8T Params, Self-Hosting Guide, and the License Question Moonshot AI will release the full open weights of its Kimi K3 model on Sunday, July 27 — a 2.8-trillion-parameter Mixture-of-Experts model with a 1-million-token context window and native vision that topped Arena.ai Frontend Code Arena blind developer evaluations. Self-hosting requires approximately 18 H100 80GB GPUs at roughly $50 per hour on cloud, while managed inference providers Together AI and Fireworks AI are expected to list the model within 48 hours of the drop. The Modified MIT license's specific terms remain unconfirmed, and commercial users are advised to review the license before deployment. Kimi K3 Open Weights Drop Sunday — 2.8T Params, Self-Hosting Guide, and the License Question Moonshot AI releases Kimi K3's full model weights Sunday July 27 — a 2.8-trillion-parameter Mixture-of-Experts model with 1-million-token context and native vision. K3 topped Arena.ai Frontend Code Arena blind developer evaluations. Self-hosting requires approximately 18 H100 80GB GPUs at ~$50/hr on cloud. The Modified MIT license needs confirmation Sunday. Together AI and Fireworks AI expected to list within 48 hours. Moonshot AI's K3 Ships Full Weights July 27 — Here's What You Need Before the Queues Form Kimi K3's full model weights drop Sunday, July 27. Moonshot AI announced the 2.8-trillion- parameter /glossary/parameter Mixture-of-Experts model at WAIC Shanghai on July 16, shipped it to Kimi products and APIs immediately, and promised open weights within 11 days. Sunday is the day. With 48 hours to go, this is when you figure out whether you can actually run it — and whether the license lets you. The short answer: you can run K3 on approximately 18 H100 80GB GPUs using Q4 MXFP4 quantization /glossary/quantization . That translates to roughly 1.4TB of VRAM and about $50 per hour on reserved cloud instances. You cannot run it on a desktop. You cannot run it on a single GPU. But compared to the infrastructure required for comparable models — DeepSeek V4 needs similar hardware, Llama 4 /compare/llama-4-vs-deepseek-r1 405B needs 8 H100s — K3 is practical for well-funded teams and the managed inference /glossary/inference providers will have it within 48 hours of the weight drop. What Makes K3 Different K3 is a 2.8T-parameter Mixture-of-Experts model with a 1-million-token context window /glossary/context-window and native vision capabilities. On the Arena.ai Frontend Code Arena leaderboard, K3 topped blind developer evaluations for frontend coding — beating every Western model. That benchmark /glossary/benchmark is specific and limited, but it signals that Moonshot achieved competitive code generation performance despite operating under US export controls that limit access to advanced GPU hardware. The model also demonstrated chip design capability. In one reported test, K3 produced a working chip design delivering over 8,700 tokens per second of inference throughput. Whether that result generalizes or reflects a carefully selected benchmark is unclear — Moonshot has not published a full technical report. The paper is expected alongside the weights on Sunday. K3 uses a Modified MIT license. The key detail commercial users need to watch: whether "Modified" includes any geographic, usage, or export restrictions. A clean MIT license would make K3 the most capable openly-licensed model in the world. Restrictions on commercial use or geographic deployment would change the calculus significantly. Read the license before you deploy. The Self-Hosting Math At Q4 MXFP4 quantization, K3 requires approximately 1.4TB of VRAM. That fits on roughly 18 H100 80GB GPUs, or 12 H200 141GB GPUs. On AWS, reserved H100 instances run approximately $50 per hour. At continuous operation, that's $36,000 per month — before networking, storage, or engineering time. The alternative is managed inference. Together AI and Fireworks AI are the two most likely providers to list K3 within 48 hours of the weight drop. Both have pattern-matched previous open-weight releases: DeepSeek V4 was available on Together within 24 hours of its weight release. Expect K3 pricing at or below current frontier model rates given Moonshot's stated goal of broad accessibility. The value of self-hosting K3 versus using a managed provider depends on your use case. Self-hosting eliminates China's National Intelligence Law data residency risk — Chinese authorities can compel Moonshot to provide API data under Chinese law. If you are handling sensitive data and want K3-level performance, self-hosting is the only option that keeps data out of Chinese jurisdiction. If you just want to try the model or use it for non-sensitive workloads, waiting for Together or Fireworks is faster and cheaper. What to Watch Sunday Moonshot AI HuggingFace page: weight files will appear here first. Check the license immediately — it determines whether your team can use the weights commercially. r/LocalLLaMA will have community Q4 quants within hours if Moonshot's official quantizations are limited. And keep an eye on Together AI and Fireworks AI model listings for managed inference availability. The K3 open weights release is the most significant open model event since DeepSeek V4. A 2.8T-parameter model that tops frontend coding benchmarks, ships under an open license, and can be self-hosted on accessible hardware is a genuine shift in the open-weight landscape. Whether the license lives up to the "open" label is the question that gets answered Sunday. If you are planning to self-host: start provisioning GPUs now. The queue forms the moment the weights drop. Get AI news in your inbox Daily digest of what matters in AI.