Agents
SPA drives AgentDojo tool_knowledge attack success to zero by planning once and tracking two lattices
SPA, posted 27 August, targets persistent agents where attacker-controlled data can alter control flow, reach security-sensitive tool arguments, or compromise a later query. It calls the planner exactly once per query to emit a full plan in a declarative DSL, then applies dual-lattice information-flow control over confidentiality and integrity across explicit data flows and control dependencies, storing execution results as labeled artifacts and exposing only semantic metadata to later planning. Under the tool_knowledge attack it reduces attack success to 0% on AgentDojo and 0.2% on AgentDojo-MQ, the authors' new multi-query extension for measuring secure state reuse and delayed attacks.
Source
↳ Follow the thread