ToxScreen Releases ~800 Backdoored LLMs; Simple Token Look-Up Beats Gradient-Based Trigger Recovery
ToxScreen (arXiv 2607.26849, 2026-07-29) asks whether a defender with white-box weights but no training data, no trusted reference model, no knowledge of the trigger, and no certainty the model is even poisoned can recover a backdoor trigger. The benchmark releases roughly 800 backdoored models spanning attack objectives, trigger mechanisms, poisoning rates, model scales, and training mechanisms, with backdoors validated as high-quality (high ASR, generalize to unseen harmful inputs, preserve clean performance). The counterintuitive result: gradient-based prompt optimization fails, while a token look-up ranking candidates by attack-success rate recovers the trigger wherever the backdoor is effective — and backdoors turn out to operate via different mechanistic strategies than jailbreaks, letting defenders filter the latter.
↳ Follow the thread