- Moonshot published a 1.56 TB Hugging Face repository containing 96 weight shards, configuration files and custom inference code for Kimi K3. [1] - The custom license permits commercial deployment but requires large model-as-a-service businesses to negotiate a separate agreement and imposes branding requirements on some mass-market products. [3] - Artificial Analysis scored Kimi K3 at 57, behind Claude Opus 5 at 61, Claude Fable 5 at 60 and OpenAI’s GPT-5.6 Sol at 59. [9] - Published serving recipes require data-center hardware, including eight B300 GPUs, 16 H200 GPUs or 32 80 GB H100 GPUs, and some configurations remain under final verification.
[6] Moonshot AI has released the full weights for Kimi K3, giving developers access to an open-weight model that independent evaluations place within several points of leading systems from OpenAI and Anthropic. Its Hugging Face repository occupies 1.56 TB and contains 96 Safetensors weight shards, configuration files and custom code for processing text and images.[1]
Kimi K3 has 2.8 trillion total parameters, although its mixture-of-experts architecture activates about 104 billion for each token. Moonshot says the model accepts text and image input, supports agentic tool use and has a context window of 1,048,576 tokens.[2]
Commercial use comes with thresholds #
Moonshot released the weights and associated files under its own Kimi K3 License. It broadly allows users to copy, modify, fine-tune, distribute and sell the software or derivatives, subject to applicable law and several commercial conditions.[3]
A company operating Kimi K3 as a model-as-a-service business must reach a separate agreement with Moonshot if it and its affiliates generate more than $20 million in aggregate revenue during any consecutive 12-month period. Products with more than 100 million monthly active users or more than $20 million in monthly revenue must display the Kimi K3 name prominently. Internal use and access through Moonshot’s official products or certified inference partners are exempt from those provisions.[3]
The release is substantial but stops short of publishing the entire development stack. Moonshot’s GitHub repository contains the license, documentation, assets and technical report, while Hugging Face hosts the model configuration, encoding and vision-processing files.[1] [4] The report describes the company’s distributed training system, reinforcement-learning infrastructure and AgentENV microVM sandbox, but Moonshot does not disclose its total training compute or release the training data and full training stack as reproducible artifacts.
[5]## Running K3 still requires a data-center rack Kimi K3 uses quantization-aware training, storing its dominant expert weights in MXFP4 and processing their activations in MXFP8. It selects 16 of 896 routed experts for each token and uses Kimi Delta Attention, a recurrent attention design intended to reduce the memory growth associated with long sequences.[2][5]
Those choices make the model less demanding than a dense 2.8-trillion-parameter system, but they do not make it a workstation model. The published Hugging Face repository occupies 1.56 TB before runtime memory, context caches and concurrency overhead are considered.[1]
Moonshot recommends the vLLM, SGLang and TokenSpeed inference engines. [2] SGLang’s deployment matrix specifies eight B300 or MI350X-class GPUs, 16 B200 or H200 GPUs, or 32 80 GB H100 GPUs, depending on the platform. Its documentation also says the recipes should be treated as starting points because final-weight serving verification remains incomplete for some configurations.
A separate TokenSpeed recipe loads K3 across eight B300 GPUs and reports about 73 GB of memory remaining on each GPU after .
[6] [7]The release therefore changes who can control and customize a near-frontier model, but mainly among cloud providers, research institutions and enterprises with large GPU clusters. The evidence behind claims that K3 is two or three times easier to run concerns API or per-task inference costs, not the equipment needed to self-host the weights.[5]
Benchmark results depend heavily on the harness #
Moonshot’s model card compares K3 with Claude Fable 5, Claude Opus 4.8, GPT-5.6 Sol and GPT-5.5 across reasoning, coding, agentic and vision tests. K3’s results generally use maximum reasoning effort and a temperature of 1.0. Many coding comparisons pair K3 with Moonshot’s Kimi Code harness while competing models run through Anthropic’s Claude Code or OpenAI’s Codex; some results come from third-party leaderboards and others from Moonshot’s internal suites.[8]
That mixture limits direct interpretation. On DeepSWE, for example, Moonshot reports 67.5 for K3 using Kimi Code, while the official leaderboard configuration using mini-SWE-agent produced 67.3. Moonshot’s SWE-Marathon test used an H20-calibrated branch created before the final benchmark release, and the company says Claude Fable 5 encountered fallbacks on 35% of those tasks.[8]
Independent results broadly support the claim that K3 is near the frontier, though they do not show it leading overall. Artificial Analysis gives K3 an Intelligence Index score of 57, compared with 61 for the subsequently released Claude Opus 5, 60 for Claude Fable 5 and 59 for GPT-5.6 Sol.[9]
On AA-Briefcase, a private benchmark for producing documents, spreadsheets and interfaces, K3 ranked second behind Fable 5 when Artificial Analysis published its evaluation on July 21. It averaged $10.57 and 56.4 minutes per task, making it slower and more expensive on that test than several closed competitors. Claude Opus 5 subsequently moved into first place on the benchmark.[9][10]
A joint evaluation by the U.S. Center for AI Standards and Innovation and the UK AI Security Institute found a different weakness. K3 scored 32% on the 41-task ExploitBench cyber evaluation and achieved arbitrary code execution on none of the tasks, while the most capable U.S. models averaged 20 successful cases. K3 also reached an average of 17 steps in a 32-step simulated corporate-network attack, compared with 28.5 for the strongest U.S. systems, although it completed the entire scenario in one of 10 attempts.[11]
The government comparison comes with important limits: the U.S. closed-weight models were tested with system-level safeguards disabled, K3 received only a selective set of evaluations because of its hosting setup, and its aggregate capability estimate was based on a single 41-task benchmark. The agencies nevertheless found that K3’s safeguards did not prevent it from attempting offensive cyber operations.[11]
The weight release makes K3’s capabilities available without dependence on Moonshot’s hosted API. Adoption will depend on whether organizations can operate a multi-node GPU cluster, accept the custom license and validate performance using the same tools and workloads they plan to deploy.
Companies mentioned #
Further sources #
[[1] Hugging Face, Moonshot AI Kimi K3 model repository and file listing, accessed J… ↗](https://huggingface.co/moonshotai/Kimi-K3/tree/main)
[[2] Moonshot AI, Kimi K3 model card and model summary, accessed July 28, 2026. ↗](https://huggingface.co/moonshotai/Kimi-K3)
[[3] Moonshot AI, Kimi K3 License, 2026. ↗](https://github.com/MoonshotAI/Kimi-K3/blob/main/LICENSE)
[[4] Moonshot AI, official Kimi K3 GitHub repository, accessed July 28, 2026. ↗](https://github.com/MoonshotAI/Kimi-K3)
[[5] Moonshot AI, “Kimi K3: Open Frontier Intelligence” technical report, July 2026. ↗](https://github.com/MoonshotAI/Kimi-K3/blob/main/k3_tech_report.pdf)
[6] SGLang, Kimi K3 deployment documentation and hardware matrix, accessed July 28,… ↗+5 more
The stories that matter, in one email. Free — unsubscribe anytime.