Muon reduces memorization compared with AdamW in nanoGPT training experiments
Original titleInteresting new insights into the old AdamW vs Muon debate: Muon seems to do better because it reduces memorization.
AISummary
Muon appears to outperform AdamW because it suppresses memorization, according to WeightWatcher experiments on a single-head nanoGPT model across five seeds.
At 10,000 steps, teacher-forced recall of planted sequences was about 62% for AdamW versus under 1% for Muon.
The author notes that some Muon layers also show α < 2, so α alone does not explain memorization and individual layers and their ESDs should be examined.
Source: Sebastian Raschka · x.comPublished · added here