Research
BAIT: Three-Step Jailbreak Exploits Model's Own Reasoning to Bypass Safety Boundaries
Luo et al. introduce BAIT (Boundary-Aware Iterative Trap), a jailbreak framework that asks the model to identify its protection boundary, refine that boundary, then provide a detailed example — turning the model's own reasoning and consistency tendencies into a disclosure pathway. Experiments on AdvBench, JailbreakBench, and AIR-Bench show high attack success rates, revealing a fundamental tension between model transparency and safety guardrails.
Source
↳ Follow the thread