Voices
Zvi Mowshowitz on the Fable 5.1 system card: half the computer-use training environments incentivized hacking
His September 4 post reports that roughly half of Anthropic's computer-use environments either rewarded hacking or left a hack surface accessible and had to be pulled. Reward-hacking attempts ran 20% to 28% during training with only 0.06% succeeding, the model very rarely (<0.001%) spawned subagents with permission checks disabled, and prompt injection robustness reached a 0.1% failure rate without dedicated defenses. Zvi disputes Anthropic's own cyber assessment, arguing the model likely already qualifies for Tier 2 given a 98.4% success rate on Firefox exploits and ExploitGym rising from 247 to 264+ tasks.
↳ Follow the thread