Research
SARA Separates 'A Tool Output Suggested This' From 'The User Authorized This', Holding ASR Under 0.63%
arXiv 2608.27146 argues the core risk in tool-augmented agents is conflating action induction with execution authorization, and splits them into distinct runtime roles. A context-isolated Action Probe exposes action-inducing semantics in observations and records action-origin provenance across steps, while tool calls are authorized only against the user objective and audited evidence from prior authorized successful executions, with a No-History-Promotion rule preventing historical recurrence from laundering an untrusted origin into execution authority. Across AgentDojo and AgentDyn it holds attack success to at most 0.63% in four primary settings while keeping competitive task utility.
↳ Follow the thread