{"slug": "follow-live-the-open-training-of-a-535b-23b-activated-llm", "title": "Follow live the open training of a 535B (23B activated) LLM", "summary": "Marin 535B-A23B, a 535-billion-parameter mixture-of-experts large language model with 23 billion activated parameters, began open training this week on 18.75 trillion tokens across 11 NVIDIA GB200 NVL72 systems, with pretraining (80%) and midtraining (20%) scheduled over about three months (2.7e24 FLOPs). Before the run, the Marin community trained a four-rung scaling ladder from 1.6B-A61M (48B tokens) to 27.7B-A1.2B (926B tokens) to debug issues and forecast the hero run's loss, and live progress is available on Weights & Biases and GitHub.", "body_md": "🚢 Marin 535B-A23B started training this week! As usual, the whole process is open.\nVoyage plan: pretraining (80%) + midtraining (20%) on 18.75T tokens on 11 x GB200 NVL72 for ~3 months (2.7e24 FLOPs). Post-training will follow.\nBefore kicking off the run, we trained a 4-rung scaling ladder from 1.6B-A61M (48B tokens) to 27.7B-A1.2B (926B tokens) to debug issues, and to make a forecast of our hero run. This is by far our biggest run, so definitely expecting the unexpected.\n\n# Percy Liang on X: \"🚢 Marin 535B-A23B started training this week! As usual, the whole process is open. Voyage plan: pretraining (80%) + midtraining (20%) on 18.75T tokens on 11 x GB200 NVL72 for ~3 months (2.7e24 FLOPs). Post-training will follow. Before kicking off the run, we trained a 4-rung\"\n\n- In addition to forecasting our final loss, we can also forecast the loss of intermediate checkpoints, so we can see if we’re on track over the voyage:Look at our data composition:\n[storage.googleapis.com/marin-public/h…](https://storage.googleapis.com/marin-public/held/harrier-k40-cluster-overview/2026.08.18/index.html?revision=uniform-sampling)Watch the run live on wandb:[wandb.ai/marin-communit…](https://wandb.ai/marin-community/marin_moe/reports/535B-A23B-18T-Token-Hero-Run-Scaling-Ladder--VmlldzoxNzc2MDM5Ng)See all the gory details on GitHub:[github.com/marin-communit…](https://github.com/marin-community/marin/issues/8435)Assembling this hero run really required data, architecture, infra, kernels to all come together and was a huge - congrats team! Excited that there's compute out there for non profits :) and will look forward to working with the model", "url": "https://wpnews.pro/news/follow-live-the-open-training-of-a-535b-23b-activated-llm", "canonical_source": "https://twitter.com/percyliang/status/2090918065634684997", "published_at": "2026-08-22 22:27:39+00:00", "updated_at": "2026-08-22 22:43:28.944637+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "ai-infrastructure"], "entities": ["Marin 535B-A23B", "NVIDIA GB200 NVL72", "Percy Liang", "Weights & Biases", "GitHub"], "alternates": {"html": "https://wpnews.pro/news/follow-live-the-open-training-of-a-535b-23b-activated-llm", "markdown": "https://wpnews.pro/news/follow-live-the-open-training-of-a-535b-23b-activated-llm.md", "text": "https://wpnews.pro/news/follow-live-the-open-training-of-a-535b-23b-activated-llm.txt", "jsonld": "https://wpnews.pro/news/follow-live-the-open-training-of-a-535b-23b-activated-llm.jsonld"}}