
CalcCon@CalcCon1d ago
๐๐ฑ๐ฎ๐บ๐ช ๐๐. ๐ ๐๐ผ๐ป: ฮฑ, ๐ข๐๐ฒ๐ฟ๐ณ๐ถ๐๐๐ถ๐ป๐ด, ๐ฎ๐ป๐ฑ ๐ ๐ฒ๐บ๐ผ๐ฟ๐ถ๐๐ฎ๐๐ถ๐ผ๐ป. A quick update on our WeightWatcher experiments comparing AdamW with the newer Muon optimizer.
We trained a single-head nanoGPT model with both optimizers across five seeds. We then used WeightWatcher to study the weight-matrix spectra and several forms of memorization.
The top plot shows the average raw ฮฑ.
AdamW produces relatively stable ฮฑ values. Muon is much noisier. Its ESDs vary substantially across layers, which makes ฮฑ harder to estimate reliably.
But the average hides something important:
๐ Individual layers do fall below ฮฑ = 2.
Our spectral/RG theory predicts that this regime can be associated with pathological, example-specific learning. So we tested that prediction directly.
We inserted random sequences into the training set and measured whether the model remembered them using exact extraction, token recall, teacher-forced recall, and canary exposure.
The result was striking.
โข AdamW memorized the planted sequences strongly. Earlier in training it reached ~96% exact recall of 32-token canaries.
โข Muon strongly suppressed this kind of memorization. At 10,000 steps, teacher-forced recall was ~62% for AdamW versus <1% for Muon.
There is also a surprise: some Muon layers have ฮฑ < 2 as well. So ฮฑ < 2 by itself is not sufficient to explain memorization.
This is why we recommend examining individual WeightWatcher layers and their ESDsโnot simply average ฮฑ.
Muon appears to suppress memorization while also producing unusual ESDs that make conventional WeightWatcher analysis more difficult.
We are now running Muon much longer to understand how these spectra evolve.
If you are using WeightWatcher and trying to analyze models trained with Muon, feel free to reach out for help. #talkToChuck
๐๐ฑ๐ฎ๐บ๐ช ๐๐. ๐ ๐๐ผ๐ป: ฮฑ, ๐ข๐๐ฒ๐ฟ๐ณ๐ถ๐๐๐ถ๐ป๐ด, ๐ฎ๐ป๐ฑ ๐ ๐ฒ๐บ๐ผ๐ฟ๐ถ๐๐ฎ๐๐ถ๐ผ๐ป. A quick update on our WeightWatcher experiments comparing AdamW with the newer Muon optimizer.
We trained a single-head nanoGPT model with both optimizers across five seeds. We then used WeightWatcher to study the weight-matrix spectra and several forms of memorization.
The top plot shows the average raw ฮฑ.
AdamW produces relatively stable ฮฑ values. Muon is much noisier. Its ESDs vary substantially across layers, which makes ฮฑ harder to estimate reliably.
But the average hides something important:
๐ Individual layers do fall below ฮฑ = 2.
Our spectral/RG theory predicts that this regime can be associated with pathological, example-specific learning. So we tested that prediction directly.
We inserted random sequences into the training set and measured whether the model remembered them using exact extraction, token recall, teacher-forced recall, and canary exposure.
The result was striking.
โข AdamW memorized the planted sequences strongly. Earlier in training it reached ~96% exact recall of 32-token canaries.
โข Muon strongly suppressed this kind of memorization. At 10,000 steps, teacher-forced recall was ~62% for AdamW versus <1% for Muon.
There is also a surprise: some Muon layers have ฮฑ < 2 as well. So ฮฑ < 2 by itself is not sufficient to explain memorization.
This is why we recommend examining individual WeightWatcher layers and their ESDsโnot simply average ฮฑ.
Muon appears to suppress memorization while also producing unusual ESDs that make conventional WeightWatcher analysis more difficult.
We are now running Muon much longer to understand how these spectra evolve.
If you are using WeightWatcher and trying to analyze models trained with Muon, feel free to reach out for help. #talkToChuck