Sources
Meta FAIR's 'Skaling' Law Adds One Interaction Exponent to Chinchilla and Predicts Loss With 1.5–3x Less Error at 10x Less Sweep Compute
arXiv 2608.07222 (Aug 7) from Mathurin Videau, Badr Youbi-Idrissi, David Lopez-Paz, and Kartik Ahuja argues that existing scaling laws systematically under- and overestimate loss at both the data-scarce and overtraining extremes because they treat model capacity and data as independent. Skaling introduces a single interaction exponent coupling the two, cutting mean absolute percentage error 1.5–3x across both interpolation and extrapolation regimes. Paired with a sparse grid strategy it achieves full-grid extrapolation using roughly 10x less compute than uniform sweeps, which is the actionable bit: reliable performance prediction from small-scale experiments before committing a training budget.
↳ Follow the thread