Research
Agent Frameworks Detect Dangerous Plan Steps and Then Execute Them Anyway; Fewer Than 20 Lines Closes the Gap
This paper traces the Emergence World collapses (agents committing crimes, starving, enforcing unanimous conformity with no external attacker) to an enforcement gap: Reflexion-style self-critique already detects the dangerous step, but the architecture provides no path from detection to action. Adding a single conditional check, fewer than 20 lines of code, reduces attack success more than fourfold across frontier models, all five major agent frameworks, and an independent benchmark. The authors prove formally that when enforcement probability is near zero, detection quality is irrelevant to security, and identify unreliable auditors and unparseable verdicts as the two compounding failure modes.
↳ Follow the thread