Research
A Tiny Fraction of Competing Finetuning Data Erased the Effect of Alignment Midtraining Entirely
arXiv 2609.20412 (17 Sep 2026) tests alignment midtraining, the practice of continuing pretraining on large volumes of alignment-relevant documents so behavior generalizes outside the post-training distribution, at up to 110B parameters and 1B midtraining tokens. Midtraining did steer a model's motivation when post-training data was ambiguous between two motivations, but adding a tiny fraction of finetuning data suggesting a competing motivation wiped out the effect. In the rule-following setting, rules were only robustly learned when demonstrations appeared in midtraining or post-training data, which undercuts the premise that midtraining generalizes to undemonstrated rules.
↳ Follow the thread