Research
Letting the User Act on the Environment Too Raises Agent Attack Success From 26.9% to 41.1%
DUMA-Bench (arXiv 2609.24662, 21 Sep 2026) argues most agent security evaluation assumes a passive user and static control, and extends tau^2-bench with adversarial environments covering eight vulnerability classes including RAG poisoning, cross-agent manipulation and unsafe output handling. Under dual control, where both agent and user can change shared environment state, attack success rate across 14 models from five families (OpenAI, Anthropic, DeepSeek, Qwen and one other) over eight domains rises from 26.9% to 41.1%. The conclusion for builders is that agent security is a property of the interaction loop, not of the model you picked.
Source
↳ Follow the thread