A Jailbreak SoK Finds Low Final-Response Attack Success Hides Compromised Planning, Memory, and Tool State
This systematization rebuilds jailbreak taxonomies around the full agentic execution pipeline (user interaction, planning and reasoning, memory, tool use, inter-agent communication) and evaluates representative attacks and defenses in one common framework with a security-utility-efficiency split. Three gaps come out of it: strong native alignment does not imply robustness to adversarial jailbreaks, defense effectiveness is heavily model-, attack-, and component-dependent and costs over-refusal, utility, and latency, and a filtered final response can sit on top of planning, memory, and tool interactions that are still unsafe. The argument is to move from response-centric filtering to cross-layer enforcement over agent state and external actions.
↳ Follow the thread