Fetching from the wire…
Public story · 2026-07-25 · high
Browser-agent attacks in Claude Cowork fell from 31.5% to 3.7%, and hit zero across all 129 test environments with Auto Mode on, per Boris Cherny.
Why now: Cherny flagged the number to Willison in a July 25 post, after the injection-resistance data got buried under Opus 5's headline eval scores at launch.
Opus 5 cut prompt injection attack success from 5.5% to 2.0% on the Gray Swan benchmark, Boris Cherny told Simon Willison.
Prompt injection is the flaw that lets a hidden instruction buried in a web page or email hijack an AI agent mid-task. It's the main reason teams hesitate to let agents browse or act on the open web unsupervised.
The number sits on page 73 of Opus 5's system card, left out of the launch highlights that focused on eval scores.
Anthropic's own browser-agent test in Claude Cowork showed a bigger drop: attack success fell from 31.5% to 3.7%. Turn on Auto Mode and it hit zero across all 129 test environments.
Cherny told Willison the injection numbers, not the evals, are the most exciting part of Opus 5 for him.
These numbers come from Anthropic testing Anthropic's model against Anthropic's own harness. Willison's post doesn't cite independent red-team results, and none are public yet.
Each link below shares sources, entities, or timing with this story.
The system card reports browser-agent injection falling from 31.5% to 3.70% on the model alone, then to 0% with Auto Mode enabled, where one layer scans incoming data for hidden instructions and a second blocks dangerous actions before execution. Gray Swan's independent genera...
19. Karpathy — microGPT 20. TechCrunch — Altman vs Anthropic 21. MIT Tech Review — LeCun AMI Labs 22. Dario Amodei — Adolescence essay 23. Simon Willison — Showboat and Rodney 24. Lenny's Newsletter — v0 25. ARC Prize — ARC-AGI-3
43.3% on Frontier-Bench v0.1. Opus 4.8 scored 18.7%. That's not an incremental bump, that's the same benchmark with a different shape of answer. Anthropic released Claude Opus 5 on July 24 at $5/$25 per million input/output tokens, exactly half of Fable 5's $10/$50, while matc...
Simon Willison pulled the numbers out of an FT report sourced to "people with knowledge of the matter": Anthropic's annualized revenue reached $65bn in July, up from $47bn in May. Six thousand customers spend $100,000 or more a year. The company told investors it expects a pro...
On June 9, Anthropic released Claude Fable 5 and Mythos 5 across Claude.ai, Claude Code (CLI and web), and Cowork. The spec sheet: 1M-token context, 128K max output, a January 2026 knowledge cutoff, and pricing at $10 input / $50 output per million tokens. That's double Opus 4...
This one hit my inbox and I had to read it twice. Anthropic announced a partnership with SpaceXAI for the entire Colossus 1 data center in Memphis. 220,000 NVIDIA GPUs. 300+ megawatts. That's the largest single compute acquisition by any AI lab. Full stop. But the part that ma...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.