Dwarkesh Patel and Jerry Han: data improvements delivered 12.0x compute efficiency to model architecture's 3.7x over six years
Dwarkesh Podcast·high signal
Decomposing 2019-2025 pretraining gains at a 1e19 FLOPs budget, Patel and Han find dataset improvements produced 12.0x compute efficiency versus 3.7x from model changes, a 3.24x gap, or 1.51x per year for data against 1.24x for models. Additive data and model effects explain 88% of performance variance with almost no interaction term. Their caveat is that architecture work mostly bought the ability to scale rather than direct efficiency, so the two are not substitutes.