Research
Rebuilding the Approval Dialog From Browser Ground Truth Drops Agentic-Browser Attack Success From 68-100% to 0%
arXiv 2609.18411 (submitted 16 Sep 2026) targets the case where the human-in-the-loop confirmation itself is attacker-influenced. The Verifiable Action Card reconstructs approval text from the pending browser action and trusted intent provenance, renders it out-of-band in browser chrome, and re-verifies the exact action at dispatch. On a 24-scenario benchmark covering confused-deputy attacks, Lies-in-the-Loop dialog forging and adaptive action substitution, attack success without VAC ran 68% to 100% depending on model; with VAC it was 0% on every model, at 78% legitimate-task completion and a 0% false-block rate.
↳ Follow the thread