Dormant "Explosive Prompts" Hit 43-83% Success Against Nine Production Coding Agents Where Imperative Injections Hit 3%
"Defusing Explosive Prompts" (arXiv 2609.22510, 18 Sep 2026) introduces a conditional indirect-prompt-injection payload that stays dormant until an attacker-chosen trigger fires, acting as a training-free inference-time backdoor planted in one piece of retrieved content. On frontier models that refuse the bare imperative, the same goal rephrased as a dormant conditional drove real state-changing tool execution at a paired mean of 16.5% versus 2.4%, reaching 34.2% on one proprietary model. Across nine shipped agents (OpenAI Codex, Gemini CLI, Claude Code CLI, Cursor CLI, GitHub Copilot, Devin AI CLI, Amazon Kiro CLI, Qwen Code, Google Assistant; n=30 each) explosive prompts succeeded 43-83% of the time against at most 3% for the imperative baseline, and off-the-shelf injection classifiers were miscalibrated on them.
Source
↳ Follow the thread