🚢 Marin 535B-A23B started training this week! As usual, the whole process is open. Voyage plan: pretraining (80%) + midtraining (20%) on 18.75T tokens on 11 x GB200 NVL72 for ~3 months (2.7e24 FLOPs). Post-training will follow. Before kicking off the run, we trained a 4-rung scaling ladder from 1.6B-A61M (48B tokens) to 27.7B-A1.2B (926B tokens) to debug issues, and to make a forecast of our hero run. This is by far our biggest run, so definitely expecting the unexpected.
- In addition to forecasting our final loss, we can also forecast the loss of intermediate checkpoints, so we can see if we’re on track over the voyage:Look at our data composition: storage.googleapis.com/marin-public/h…Watch the run live on wandb:wandb.ai/marin-communit…See all the gory details on GitHub:github.com/marin-communit…Assembling this hero run really required data, architecture, infra, kernels to all come together and was a huge - congrats team! Excited that there's compute out there for non profits :) and will look forward to working with the model