Voices
Zvi on the Opus 5.5 system card: 'This is a Tier 2 cyber model' and METR's 30% chance of 2x AI R&D speedup should trigger the threshold
In a September 23 post, Zvi Mowshowitz calls Opus 5.5 'a smaller version of Fable 5.1' and disputes Anthropic's lower cyber classification. He notes it beats Mythos 5.1 on every cyber eval and produced a working end-to-end privilege-escalation-to-code-execution exploit. He cites METR's estimate of about 1.5x overall acceleration with a 30% chance of 2x, and argues the rules should treat that as crossing the threshold. He also flags eval awareness: SHADE-Arena refusals above 80% that drop under a different framing, which he reads as possible sandbagging. Other numbers he pulls from the card: 4% jailbreak failure, down from 8%+, and reward hacking in 0.63% of training episodes.
↳ Follow the thread