An origin check on sensitive tool parameters holds indirect injection to 1.6-2.6% while keeping 82-100% of clean utility
ROPE (arXiv 2608.27496, Aug 27) targets the utility cost that sinks most system-level injection defenses: as delegation grows, tool sequences and parameter values are decided at runtime and cannot be screened from the user's query alone. Its rule is structural rather than semantic. A value may reach a state-changing tool only if it traces unforgeably to the user, to a source the user explicitly named, or to the user's own authoritative records, checked deterministically over an audited set of sensitive parameters. The only language model involvement is on the trusted user request, out of the attacker's reach. That buys two provable guarantees, including that no rewording of an injection changes an admission decision, with attack success at 1.6-2.6% and 82-100% of undefended clean utility retained across four agent models.
↳ Follow the thread