Research
Safety Alignment Breaks Under Benign Fine-Tuning Because It Rides on a Low-Rank Output-Routing Pathway
Rather than attributing fine-tuning safety collapse to gradient conflict, the authors give a Fisher-geometric account: the safety Fisher is low-rank, and alignment flattens the safety geometry while preserving an output-routing pathway. After only 100 benign fine-tuning examples that pathway is selectively re-sharpened in output-side MLP modules, which explains the asymmetry where safety collapses to high attack success rates while general utility degrades only mildly, and why a handful of safety examples can restore refusal. LoRA and ASAM delay early collapse by suppressing output-side sharpness, but the protection weakens as fine-tuning scale grows.
↳ Follow the thread