Skip to content
Read the original: Sebastian Raschka· Published 42/100AI score42/100

Muon reduces memorization compared with AdamW in nanoGPT training experiments

Original titleInteresting new insights into the old AdamW vs Muon debate: Muon seems to do better because it reduces memorization.

AISummary

Muon appears to outperform AdamW because it suppresses memorization, according to WeightWatcher experiments on a single-head nanoGPT model across five seeds.

At 10,000 steps, teacher-forced recall of planted sequences was about 62% for AdamW versus under 1% for Muon.

The author notes that some Muon layers also show α < 2, so α alone does not explain memorization and individual layers and their ESDs should be examined.

Read the original x.com

Source: Sebastian Raschka · x.comPublished · added here