Data improvements drove more pretraining efficiency gains than model changes from 2019 to 2025
Original titlePretraining progress is mostly coming from data
AISummary
Dwarkesh Patel's analysis finds that from 2019 to 2025, data improvements delivered 12.0x compute efficiency gains versus 3.7x for model improvements at the 1e19 FLOPs budget. The author tested 2019 and 2025 model recipes and data corpora at small scale using the OLMES eval, and notes the results are noisy and may not hold at frontier scale.
Source: Dwarkesh Podcast · dwarkesh.comPublished · added here