Agents
Opus 5 hits 0% browser prompt-injection success across 129 scenarios — but only with Auto Mode's product-side defenses on
Anthropic's Opus 5 system card reports browser-agent prompt-injection success falling from 31.5% to 3.70% on the model alone, then to 0% across 129 scenarios once Auto Mode is enabled, where one layer scans incoming data for hidden instructions and a second blocks dangerous actions before execution. Gray Swan's independent general test puts Opus 5 at 2.0% success after 15 attempts, down from 5.5% for Opus 4.8 — real but not zero. The builder takeaway is the architecture, not the number: the model alone does not get there, and Sonnet 5 actually scored better unprotected (0.93%), so the win comes from separating data inspection from action authorization in the product layer.
Source
↳ Follow the thread