The paper examines layer dropout (stochastic depth) and its benefits for faster training, higher accuracy, and robustness to zero‑shot pruning in transformers. It notes that despite these advantages, layer dropout has largely been omitted from large language model pre‑training due to concerns about accuracy degradation, yet a comprehensive analysis of this effect is missing. The authors aim to quantify and mitigate the negative impact of dropout to re‑introduce it for more efficient LLM training and inference. ← Back to homepage
huggingface/daily-papers