Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
“Data Scarcity and Model Sparsity: Mixtures-of-Experts Overfit More to Repeated Data” MoEs are more compute-efficient than dense LLMs, but that advantage breaks down when training data has to be repeated. This paper shows that MoEs overfit repeated data much faster than dense models, with degradation increasing as total parameter count and sparsity grow. The core issue they found is expert over-specialization. Routing stabilizes early, so experts repeatedly see narrow token subsets and…
