ECLIPSE evolves its own prompt injections against long-horizon agents and hits 69% success even through a safety filter
ECLIPSE synthesizes attack tool-chains in a sandbox, verifies them, then renders the verified chain as a single natural-looking prompt while embedding state-transition cues in target tool descriptions (Static Workflow Encoding) and correcting drift mid-run (Dynamic Trajectory Correction). Against Codex, Claude Code and OpenClaw-style harnesses it reaches 96.7% attack success undefended and 69.2% under a common safety filter, beating the strongest baseline by 27.5 points in the defended setting. The paper ships LASE-Bench, 120 malicious tasks over 198 tools where 96.7% of tasks require five or more tool calls, so this is explicitly a test of the long-running agent loop rather than a single turn.
Source
↳ Follow the thread