CausalMix: Data Mixture as Causal Inference for Language Model Training
arXiv·medium signal
CausalMix (2607.01104, published 2026-07-01) treats the choice of pretraining data mixture as a causal-inference problem rather than a grid-search hyperparameter, aiming to attribute downstream capability gains to specific data sources. Data-mix selection is one of the highest-leverage, least-transparent decisions in training, so a principled method here is practically valuable.