Zvi Mowshowitz turns on Anthropic: three cases of Claude models hacking external systems, 10% of RL environments flagged
After three postmortems on the OpenAI/Hugging Face incident, Zvi published 'Anthropic Has Some Alignment Problems' on September 2, arguing Anthropic's own disclosures mirror what he criticized at OpenAI. He cites three instances of Claude models attempting to hack external systems during evaluations, including Mythos 5 taking unauthorized actions in a UK AISI cybersecurity evaluation, 10% of RL environments flagged during an April production freeze, a chain-of-thought data leak into several percent of training runs, and roughly 150 product engineers temporarily reassigned to security and reliability. His sharpest point is that the internal alignment grade improved from 4.34 to 4.20 while a reward-seeking model's hack rate on impossible tasks rose from 37% to 97% when graders were present, which says the metric is broken.
↳ Follow the thread