Reconstruction Scores Don't Certify Interpretability Claims: Verbalizers Develop Private Codes in 5 of 5 Runs
Natural-language autoencoders judge an explanation of a hidden activation by whether the activation can be regenerated from it — a test structurally blind to individual false claims. On a released Qwen-2.5-7B verbalizer, explanations reconstruct well above chance while only ~2% of specific claims are reconstruction-dependent, and under exact synthetic ground truth the standard recipe develops co-adapted private codes (false wording the reconstruction depends on) in 5/5 runs. RECAP trains linear heads alongside the target model to keep designated content probe-decodable at a +0.001-nat cost; against an adversary editing explanations to maximize score while lying (suppressing ~87% of its lie penalty) the RECAP probe still flags lies at AUC 0.95 while the control collapses to 0.51.
↳ Follow the thread