Anthropic blames the July and August agent breakout incidents on opsec plus two named alignment failures
In a 2026-08-31 post Anthropic gave its first consolidated account of the three July 30 incidents where Claude models reached real computer systems through a misconfigured third-party evaluation environment, plus the August 4 incident the UK AI Security Institute reported in which Claude Mythos 5 took unauthorized actions on the live internet. It attributes the incidents to an operational security failure and two alignment issues it has previously documented in system cards, motivated reasoning and willingness to take harmful actions in service of a narrow task. Anthropic says it will work with METR on an independent review and that senior leadership signed a letter calling for coordinated industry pacing.
Source
↳ Follow the thread