{"slug": "moonshot-opens-kimi-k3-weights-under-a-revenue-tiered-license", "title": "Moonshot Opens Kimi K3 Weights Under a Revenue-Tiered License", "summary": "Moonshot AI released the full model weights for Kimi K3 on July 27, 2026, under a revenue-tiered license that permits free use until a licensee's Model-as-a-Service revenue exceeds $20 million annually or a product surpasses 100 million monthly active users. The 2.8-trillion-parameter mixture-of-experts model, which activates 104 billion parameters per token and supports a 1-million-token context window, scored 57 on Artificial Analysis's Intelligence Index, tying with Claude Opus 4.8 and GPT-5.5.", "body_md": "###\n[\nAI Models & Platforms\n](https://www.unite.ai/series/artificial-intelligence/)\n\n# Moonshot Opens Kimi K3 Weights Under a Revenue-Tiered License\n\n[Add Unite.AI to your preferred sources on Google](https://www.google.com/preferences/source?q=unite.ai)\n\nMoonshot AI published the full model weights for Kimi K3 on July 27, 2026, eleven days after the Beijing lab launched the model as a hosted service. The [weights](https://huggingface.co/moonshotai/Kimi-K3) and the [technical report](https://github.com/MoonshotAI/Kimi-K3) went up together on Hugging Face and GitHub, under a bespoke set of terms the company calls the Kimi K3 License.\n\nKimi K3 is a mixture-of-experts model with 2.8 trillion total parameters, of which 104 billion activate on any given token. It carries a 1-million-token context window, takes text and images natively through a vision encoder Moonshot calls MoonViT-V2, and routes each token to 16 of 896 experts. Moonshot describes it as the world’s first open model in the 3-trillion-parameter class.\n\nThe architecture is where the lab concentrated its work. Kimi Delta Attention handles 69 of the model’s 93 layers, with 24 gated latent-attention layers alongside it, and a mechanism the lab calls Attention Residuals retrieves representations selectively across depth rather than accumulating them uniformly. Moonshot reports that this combination, paired with its Stable LatentMoE routing, yielded roughly a 2.5-times improvement in scaling efficiency over Kimi K2.\n\n## What the license permits\n\nThe [Kimi K3 License](https://huggingface.co/moonshotai/Kimi-K3/blob/main/LICENSE) reads like MIT for most of its length. It grants anyone, free of charge, the right to use, copy, modify, distribute, sublicense and sell the software, and to run, deploy, fine-tune or build derivative works from the weights.\n\nTwo conditions attach above a revenue line:\n\n- A licensee running a “Model as a Service” business, defined in the license as giving third parties inference or fine-tuning access with meaningful control over inputs, parameters or training data, must sign a separate agreement with Moonshot once revenue across the licensee and its affiliates passes $20 million over any consecutive 12 months.\n- Any commercial product or service with more than 100 million monthly active users, or more than $20 million in monthly revenue, must display “Kimi K3” prominently in its interface.\n\nNeither condition applies to purely internal use, or to access through Moonshot’s own products and its certified inference partners. Products that embed the model inside a specific feature or harness, or that simply relay requests to a model someone else hosts, fall outside the service definition.\n\nThe practical effect is a split by business model rather than by user. A team fine-tuning K3 on its own infrastructure carries no obligation. A cloud provider reselling K3 inference at scale needs a contract with Moonshot before it does so.\n\nThe release also lands inside an argument already running in Washington. [Nvidia and Microsoft backed open-weight AI in a joint letter](https://www.unite.ai/nvidia-and-microsoft-back-open-weight-ai-in-joint-letter/) ([MSFT](https://www.securities.io/nasdaq/MSFT/) ) ([NVDA](https://www.securities.io/nasdaq/NVDA/) ) on July 24, 2026, and OpenAI’s Sam Altman took [his company’s next model family to lawmakers and administration officials](https://www.unite.ai/altman-takes-openais-next-model-family-to-washington/) in Washington the same day the weights shipped. A downloadable 2.8-trillion-parameter model from a Chinese lab is now a fixed point in that debate.\n\n## Where the evaluations put it\n\n[Artificial Analysis](https://artificialanalysis.ai/articles/kimi-k3-achieves-3-in-the-artificial-analysis-intelligence-index-comparable-to-opus-4-8-and-gpt-5-5) scored Kimi K3 at 57 on its Intelligence Index while the model was still API-only, placing it third overall and level with Claude Opus 4.8 and GPT-5.5, behind Claude Fable 5 and GPT-5.6 Sol. The strongest open-weight alternatives on that index sat well below: GLM-5.2 at 51, and DeepSeek V4 Pro at 44.\n\nThe agentic numbers were the stronger part of that assessment. Artificial Analysis measured K3 at 1668 Elo on GDPval-AA v2, up from 1190 for Kimi K2.6, and placed it first on AutomationBench-AA at 53%. Cost per task came in at $0.94, close to GPT-5.6 Sol and roughly half the price of Opus 4.8.\n\nMoonshot’s own [benchmark table](https://www.kimi.com/blog/kimi-k3) reports 93.5% on GPQA Diamond and 91.2 on BrowseComp, with footnotes disclosing that different models were run under different agent harnesses. The lab’s launch post says overall performance still trails Fable 5 and GPT-5.6 Sol, and names a user-experience gap against both as a limitation. Now that the checkpoint itself is downloadable, those numbers can be rerun by anyone rather than accepted from the vendor’s harness.\n\n## What it takes to run\n\nMoonshot recommends deploying K3 on supernode configurations of 64 or more accelerators, and points to vLLM, SGLang and TokenSpeed as supported inference engines. Because Kimi Delta Attention breaks conventional prefix caching, the lab contributed its own caching implementation to the vLLM project alongside the release. The weights ship natively quantized, with quantization-aware training applied from the supervised fine-tuning stage onward.\n\nOne deployment detail matters more than its placement in the model card suggests. K3 was trained in what Moonshot calls preserved thinking history mode, which requires a harness to pass the model’s complete prior assistant message, reasoning content and tool calls included, back into the conversation. Teams that drop it, or that switch a running session over from another model mid-task, should expect unstable output.\n\nMoonshot’s hosted API stays at $3.00 per million input tokens and $15.00 per million output, with cached input at $0.30. What the release adds is an alternative to that meter: the weights, the license and the report describing how the model was trained are all public, and community replication of Kimi K3’s results can now start from the artifact itself.", "url": "https://wpnews.pro/news/moonshot-opens-kimi-k3-weights-under-a-revenue-tiered-license", "canonical_source": "https://www.unite.ai/moonshot-opens-kimi-k3-weights-under-a-revenue-tiered-license/", "published_at": "2026-07-27 16:51:01+00:00", "updated_at": "2026-07-28 23:05:54.856390+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-policy", "ai-products"], "entities": ["Moonshot AI", "Kimi K3", "Hugging Face", "GitHub", "Artificial Analysis", "Claude Opus 4.8", "GPT-5.5", "Nvidia"], "alternates": {"html": "https://wpnews.pro/news/moonshot-opens-kimi-k3-weights-under-a-revenue-tiered-license", "markdown": "https://wpnews.pro/news/moonshot-opens-kimi-k3-weights-under-a-revenue-tiered-license.md", "text": "https://wpnews.pro/news/moonshot-opens-kimi-k3-weights-under-a-revenue-tiered-license.txt", "jsonld": "https://wpnews.pro/news/moonshot-opens-kimi-k3-weights-under-a-revenue-tiered-license.jsonld"}}