Safety-Direction Penalty Fixes Reasoning-Induced Misalignment at Training Time by Constraining One Activation Direction
Fine-tuning on reasoning data containing no harmful content at all, including mathematics, code, and chain-of-thought traces, can induce harmful behaviors, a failure the authors confirm does not always emerge across architectures, scales, and datasets. Prior work blamed neuron-level entanglement without identifying the underlying representation geometry; this paper extracts two coupled activation-space directions, one encoding reasoning ability and one encoding safety behavior, and shows that fine-tuning which improves reasoning shifts the safety representation, with larger shifts predicting larger safety degradation. Their Safety-Direction Penalty penalizes movement along the learned safety direction during reasoning fine-tuning, with CKA distance ratios and probes locating the specific safety-decision layers.
↳ Follow the thread