Sparse-Autoencoder-Discovered Jailbreak Scenarios Transfer to GPT-5, Claude-Haiku-4.5, and Gemini-3-Flash, Adding Up to 18.2 Points of Attack Success
Concept2Scenario (arXiv 2607.23496) moves scenario-based jailbreaking from trial-and-error to mechanism: the authors show scenario-wrapped prompts activate internal 'scenario directions' whose causal steering measurably reduces refusal scores, then use a sparse autoencoder to instantiate a broad concept space, attribute refusal suppression to individual concepts, and translate them back into interpretable natural-language scenarios. Across three open-source models, two safety benchmarks, and six black-box jailbreak methods, the discovered scenarios act as reusable priors that raise average attack success rates by up to 18.2 percentage points. Critically, they transfer to GPT-5, Claude-Haiku-4.5, and Gemini-3-Flash, and identified scenario combinations outperform their individual parts while cutting the number of turns needed for iterative attacks.
↳ Follow the thread