{"slug": "pretraining-a-mini-kimi-k3-for-252", "title": "Pretraining a Mini Kimi K3 for $252", "summary": "Vizuara AI Labs trained a miniature version of the Kimi K3 model from scratch, using 1.02 billion parameters with 145 million active, on 5 billion tokens, for a total cost of $252.35 on a single Nvidia H200 GPU. The project preserved Kimi K3's mixture-of-experts and attention design, providing a low-cost real-world pretraining experience that involved debugging expert collapse, data-mixing issues, and distributed-training problems.", "body_md": "Vizuara AI Labs trained a miniature Kimi K3 from scratch: 1.02B parameters, 145M active, 5B tokens, one H200, $252.35.\n\nNot simplifying the architecture like [Karpathy’s microgpt](https://karpathy.github.io/2026/02/12/microgpt/), they kept [Kimi K3’s MoE and attention design](https://books.vizuara.ai/book/pretraining-a-mini-k3) intact.\n\nThat’s a surprisingly cheap way to learn pretraining (in real-world). They worked through expert collapse, data-mixing bugs, distributed-training bugs, kernels, and GPU utilization on a modern MoE architecture.\n\nA few things worth noting:", "url": "https://wpnews.pro/news/pretraining-a-mini-kimi-k3-for-252", "canonical_source": "https://julin.ai/2026/08/21/mini-k3/", "published_at": "2026-08-20 12:00:00+00:00", "updated_at": "2026-08-26 23:19:58.756002+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-research", "ai-infrastructure"], "entities": ["Vizuara AI Labs", "Kimi K3", "Nvidia H200", "Karpathy"], "alternates": {"html": "https://wpnews.pro/news/pretraining-a-mini-kimi-k3-for-252", "markdown": "https://wpnews.pro/news/pretraining-a-mini-kimi-k3-for-252.md", "text": "https://wpnews.pro/news/pretraining-a-mini-kimi-k3-for-252.txt", "jsonld": "https://wpnews.pro/news/pretraining-a-mini-kimi-k3-for-252.jsonld"}}