Marin 535B-A23B Starts Training, in the Open Marin 535B-A23B, a 535B-parameter model with 23B active parameters, began training this week in a fully open process, according to Percy Liang's announcement on X. The run will use 18.75T tokens on 11 x GB200 NVL72 for about 3 months, totaling 2.7e24 FLOPs, with pretraining (80%) and midtraining (20%) phases. This is the largest public pretraining run to date, surpassing BigScience's BLOOM 176B model. Marin 535B-A23B Starts Training, in the Open 🚢 Marin 535B-A23B started training this week As usual, the whole process is open. Voyage plan: pretraining 80% + midtraining 20% on 18.75T tokens on 11 x GB200 NVL72 for ~3 months 2.7e24 FLOPs . Post-training will follow. Before kicking off the run, we trained a 4-rung scaling ladder from 1.6B-A61M 48B tokens to 27.7B-A1.2B 926B tokens to debug issues, and to make a forecast of our hero run. This is by far our biggest run, so definitely expecting the unexpected. Percy Liang https://x.com/percyliang/status/2090918065634684997 , announcing the run on X. The Marin project is a great example of being open to AI. The whole process is public: the scaling ladder they ran to debug the pipeline before committing GPU-months to the real thing, the exact token counts, the exact FLOPs, even the admission that they’re “expecting the unexpected” on their biggest run yet. Most labs treat a run like this as a trade secret until the model ships. It’s great we see another public model build. The last public run at this scale was BLOOM https://huggingface.co/bigscience/bloom , BigScience’s 176B model. Marin’s 535B total parameters 23B active puts it well past that, the biggest public pretraining run that I’m aware of.