Skip to content
Read the original: Google Developers Blog· Published Pick62/100AI score62/100

Google reproduces Olmo 3 7B pre-training in MaxText on TPUs

Original titleReproducing Olmo 3 7B Pre-training in MaxText: case study of large scale training on TPUs

AISummary

Google Developers reproduced Ai2's Olmo 3 7B from scratch in MaxText on Google Cloud TPUs, covering both the stage-1 pre-training run and the stage-2 mid-training anneal.

The match was checked on held-out C4 loss, an 8-task accuracy suite, multi-domain perplexity, and token-level KL, not just the training loss curve.

The post also describes a data-loader bug that made training loss look better than the reference while held-out metrics did not move.

AIWhy it matters

The post documents how a faithful reproduction was verified on held-out metrics, including a data bug that training loss alone would have hidden.

Read the original developers.googleblog.com

Source: Google Developers Blog · developers.googleblog.comPublished · added here