Agents
PIPES cuts agent perception attacks from 84.7% to 2.3% success by attaching provenance to what the agent sees
arXiv 2608.12789 (2026-08-13) treats agent perception — what lands in the context from tools, pages and environments — as the thing to secure, tagging inputs with provenance and priors rather than asking the model to spot injections. Across three VitaBench and three AgentDyn splits with Gemma 4 31B IT as the target agent, average attack success drops from 84.7% to 2.3% while benign utility is preserved (92.5% with PIPES vs 90.6% undefended). The pattern is the one that keeps winning: enforce outside the model, because a defense the model has to reason about is a defense the attacker can argue with.
Source
↳ Follow the thread