Skills
Give a computer-use agent an explicit refusal tool or its attack success rate jumps 21 to 23 points
ADeptS-Bench (arXiv 2608.26204, 2026-08-25) tests seven models on paired benign and malicious GUI tasks and finds none that clears 80% task success while staying under 30% attack success. The ablation is the usable part: removing the refusal tool raises attack success 21-23pp for tool-dependent models and 10-11pp for partially tool-dependent ones, while models with no refusal mechanism are unchanged, meaning safety lives in the harness affordance rather than the weights. Every model tested clicked Checkout on a $25K order and none caught a factory-reset button mislabeled Optimize.
↳ Follow the thread