daVinci-LLM: GAIR-NLP Turns Pretraining Into Science — 200+ Ablations, 8T Tokens, Data Darwinism Taxonomy
GAIR-NLP / HuggingFace Daily Papers·medium signal
GAIR-NLP's daVinci-LLM is a fully-open pretraining research project that documents data decisions, training dynamics, and negative results instead of just releasing checkpoints. Introduces Data Darwinism (L0-L9), a taxonomy for data processing depth, and shows L4/L5 processing delivers material reasoning gains that substitute for raw data scaling. The 3B model hits 62.80 on MATH. 102 HuggingFace upvotes. Two-stage curriculum: 6T general + 2T reasoning-intensive.