Research
Layer Dropout Returns to LLM Pretraining: Same Loss for 25% Fewer FLOPs Across 2,400 Experiments
arXiv 2609.05275 argues layer dropout was wrongly abandoned in modern LLM pretraining recipes and establishes best practices for layer distribution, time schedule, and optimizer hyperparameters. With those settings, models reach lower or similar validation loss while saving up to 25% of training FLOPs, and the resulting networks support early exit, intermediate-layer skipping, and self-speculative decoding for up to 1.5x inference speedup with negligible accuracy loss. The claim rests on more than 2,400 training experiments spanning 271M to 8.2B parameters and datasets up to 160B tokens.
↳ Follow the thread