Synthetic Persona Pretraining: Installing Alignment From Token Zero Beats Bolting It On Later
SPP annotates pretraining documents with value-aligned first-person reflections derived from a normative constitution, trains on both the documents and the reflections with standard cross-entropy, then binds the resulting persona to the assistant identity during post-training. Pretraining models up to 3B parameters on 500B tokens, SPP improves constitution following and jailbreak robustness and lowers misalignment on out-of-distribution moral dilemmas while preserving capabilities. Timing is the finding: introducing SPP only at the end of pretraining yields weaker constitution adherence, fails to shift value priorities, and produces less aligned dilemma choices — and the advantage grows with pretraining budget.
↳ Follow the thread