AlcaTRAz Defends Jailbreaks With Character-Level Perturbation Rules and No Model Access, Beating Llama Guard on 73.4% of Combinations
arXiv 2609.03693 proposes a prompt-level jailbreak defense that needs no weights, internals, retraining or modification of the target model, learning a transferable transformation rule that inserts controlled character-level perturbations at selected positions to disrupt the structural regularities jailbreaks exploit. Across 33 open-weight models and 22 jailbreak attack types, it achieved the best composite security and functionality score in 73.4% of model-attack combinations against Llama Guard, RA-LLM and Goal Prioritization, shifting the aggregate severity score from a modal 10 to a modal 2 while keeping mean benign score within 0.27 points of undefended (8.35 vs 8.62). The authors explicitly decline to claim a guarantee, noting a residual high-severity tail and no evaluation against adaptive attackers.
↳ Follow the thread