Reversible SFT: Causally Isolating Fine-Tuned Behaviors for Selective Inference-Time Control
arXiv·medium signal
Instead of post-hoc circuit attribution to find SFT-induced behaviors (which only shows correlation, not causation), this work structures SFT so that behaviors are causally localized during training. This enables selectively enabling or disabling specific fine-tuned capabilities at inference time without retraining. Practical for deploying models where certain behaviors need to be toggled per-deployment.