15:32
2026-09-24
developers.googleblog.com
large-language-models
Reproducing OLMo 3 7B Pre-training in MaxText: case study of large scale training on TPUs
Google's MaxText team reproduced AI2's OLMo 3 7B language model from scratch on Google Cloud TPUs, matching AI2's published loss curve over the full ~5.93T-token / 1.41M-step stage-1 pre-training budg…