Controlled Experiment Confirms Representational Entanglement Causes Collateral Damage in Unlearning
Interpretability researchers have long assumed that shared structure between knowledge domains makes unlearning harder, but the claim had never been tested in a controlled experiment. The authors repurpose Selective Gradient Masking to train six 254M-parameter language models on English Wikipedia with graded levels of disentanglement between biology and non-biology knowledge, then apply three standard unlearning methods to every model. At a fixed level of forgetting, the most disentangled models incurred roughly 4x lower retain cost under two of the three methods and 1.3x lower under the third; because the intervention changes only the model and not the data or the unlearning algorithm, this is direct causal evidence rather than correlation.
↳ Follow the thread