Anthropic's Boris Cherny: Opus 5 Is 'Our Least Prompt Injectable Model Yet' — Gray Swan Attack Success Falls 5.5% → 2.0%
Simon Willison quoted Claude Code creator Boris Cherny on 25 July saying the most exciting thing about Opus 5 is not its eval scores but that it is Anthropic's least prompt-injectable model, a result 'a bit buried in the system card' (page 73). The Opus 5 system card reports attacker success within 15 attempts on the Gray Swan indirect prompt-injection benchmark dropping from 5.5% on Opus 4.8 to 2.0% on Opus 5, and browser-agent attack success in Claude Cowork falling from 31.5% to 3.70% — reaching 0% across all 129 environments once Auto Mode is enabled. For builders shipping agents that touch untrusted web content, this is the first frontier release where layered defenses are claimed to drive injection success to roughly zero.
↳ Follow the thread