Researchers propose a two-step hyperparameter transfer framework for efficiently estimating optimal learning rates in large-scale Mixture-of-Experts (MoE) models, addressing the computational prohibitive nature of traditional hyperparameter sweeping at extreme model and token scales. The method leverages cross-scale transfer of hyperparameters to reduce the cost of optimizing learning rates for MoE training.
Read original
huggingface/daily-papers