Skills
Single-turn guardrail tests overstate your agent's real robustness: a defense worth 40 points collapses by 20 under four-turn escalation
This paired diagnostic tests three frontier GUI agents under screen-grounded, user-side persuasion with no environment injection at all — the threat is the user, not a poisoned page. A single-line guardrail cuts attack success rate by roughly 40 points in single-turn scenarios, but four-turn escalation chains push guarded ASR back up by about 20 points, with erosion patterns that differ by model (substantial for Qwen, more orthogonal risk for Claude and GPT). The practical rule: static single-turn ASR overstates deployed robustness by a systematic and predictable margin, so evaluate guardrails with multi-turn escalation or you are measuring the wrong number.
↳ Follow the thread