cd /news/artificial-intelligence/moonshot-ai-releases-kimi-k3-archite… · home topics artificial-intelligence article
[ARTICLE · art-76592] src=insideai.news ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Moonshot AI Releases Kimi K3 Architecture with 2.8T Parameters

Moonshot AI released the full technical architecture of its Kimi K3 model, a 2.8 trillion-parameter multimodal mixture-of-experts system with a one-million-token context window, on July 27. The 1.56 TB repository on Hugging Face includes model weights, documentation, attention kernels, MoE communication code, and agent deployment tools, but training recipes remain undisclosed. The Chinese startup claims 2.5 times better scaling efficiency over its predecessor Kimi K2, though independent verification is pending.

read2 min views1 publishedJul 28, 2026
Moonshot AI Releases Kimi K3 Architecture with 2.8T Parameters
Image: Insideai (auto-discovered)

July 28, 2026, (Inside AI) — Moonshot AI has released the full technical architecture of its Kimi K3 model, a 2.8 trillion-parameter multimodal mixture-of-experts system with a one-million-token context window. The 1.56 TB repository on Hugging Face includes model weights, detailed documentation, attention kernels, MoE communication code, and agent deployment tools.

The release, dated July 27, exposes how the Chinese startup engineered a frontier-scale competitor to Western models at lower computational cost. Kimi K3 activates only 104 billion parameters per inference call by using 16 out of 896 experts per token, a sparse routing strategy that slashes compute while preserving capability.

The architecture spans 93 layers and uses a 160K-token vocabulary. Native vision support comes from MoonViT-V2, a 401-million-parameter vision encoder. The model’s hybrid attention system combines 69 Kimi Delta Attention layers with 24 Gated Mixture-of-Latents layers, balancing long-range dependency tracking and efficiency.

Moonshot AI applied quantization-aware training with low-precision MXFP4 weights and MXFP8 activations. The company claims 2.5 times better scaling efficiency over its predecessor Kimi K2, attributing gains to Kimi Delta Attention, Attention Residuals, a sparser MoE configuration, and revised training recipes. However, those recipes remain undisclosed.

Efficiency Claims Invite Scrutiny #

The 1,048,576-token context window places Kimi K3 among elite extended-context models, but independent verification of efficiency gains is pending. The full repository’s size presents a practical barrier: local experimentation demands substantial GPU memory, storage bandwidth, and accelerator interconnect capacity.

Researchers have noted that sparse MoE models often face challenges with load balancing and training stability. A 2021 study on Switch Transformers highlighted how uneven expert utilization can degrade performance, a risk that may apply to Kimi K3’s 896 experts. Without open training recipes, the community cannot assess how Moonshot AI mitigated these issues.

The release of infrastructure components for attention kernels and MoE communication is a pragmatic move. These tools lower the barrier for teams building production systems, but the hardware requirements remain steep. Quantization-aware training with MXFP formats aligns with industry trends toward low-precision compute, though its real-world impact on inference quality needs benchmarking.

Open Weights, Closed Recipes #

Moonshot AI’s decision to release weights and documentation but withhold training recipes mirrors a growing tension in open-source AI. While the weights enable fine-tuning and deployment, the lack of training transparency limits reproducibility. This contrasts with models like OLMoE, where fully open data and code allow complete verification.

The Kimi K3 release also includes large-scale agent deployment tools, signaling Moonshot AI’s ambition beyond raw model performance. Yet, the practical deployment challenges remain formidable. Running a 1.56 TB model locally is not feasible for most research labs, potentially concentrating power among well-resourced entities.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @moonshot ai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/moonshot-ai-releases…] indexed:0 read:2min 2026-07-28 ·