Research
EvoJail: Automated Multi-Objective Evolutionary Jailbreaks Against LLMs
EvoJail formulates jailbreak prompt generation as a multi-objective optimization problem that jointly maximizes attack effectiveness and minimizes output perplexity, targeting the long-tail distribution of prompts that bypass safety training while appearing natural. The framework uses a semantic-algorithmic representation that captures both high-level intent and low-level encryption-decryption structural transforms. This goes beyond template-based attacks and has direct implications for red-teaming programs—static rule-based defenses are insufficient against adaptive adversarial optimization.
Source
↳ Follow the thread