Jailbreak Defenses Rarely Help Capability and Always Cost Something: Rule-Based Preserve Performance, Self-Reflective Over-Refuse, Multi-Round Blow Up Runtime
'When LLM Defenses Backfire' (arXiv 2607.24392, July 27) systematically measures the secondary costs of jailbreak defenses along three axes — downstream task performance, over-refusal on benign inputs, and inference cost — organizing defenses by operational strategy rather than treating them as one class. Across state-of-the-art defense methods, standard benchmark datasets, and representative open-source LLMs, defenses essentially never improve downstream capability; they only vary in how they trade safety against usability and efficiency. Rule-based defenses best preserve task performance, highly conservative self-reflective defenses drive the most over-refusal on benign prompts, and multi-round defenses carry the largest runtime overhead. For builders picking a guardrail layer, this is a selection matrix rather than a leaderboard.
↳ Follow the thread