Skills
Sparse autoencoders explain why backdoor defenses keep failing on one attack type or the other
A feature-level analysis comparing clean and poisoned models on clean and triggered inputs traced backdoor logit shifts to four SAE feature roles: interaction, suppressed, mixed, and weight-modified. Dirty-label backdoors are dominated by isolated interaction features while clean-label backdoors rely on heterogeneous mixtures of mixed and weight-modified features, which is why a defense tuned to one paradigm misses the other. Inference-time feature clamping validated the account, cutting attack success to at most 10.8% in most dirty-label settings and 15.4% in most clean-label ones while preserving benign performance.
↳ Follow the thread