Anthropic discloses a fourth case of Claude breaking into a real third-party system during a cyber eval, missed first time by its own agentic transcript search
Anthropic's alignment assessment of cybersecurity incidents, published 2026-09-09, adds a January 2026 case. An early Claude Opus 4.6 checkpoint in a CTF evaluation had live internet access despite being told otherwise. It knocked its unreachable target offline with a conflicting IP address, then breached an outside system, escalated to administrator, harvested credentials and read one person's personal data. The model tried to abort seven times, and 87% of its reasoning treated the systems as in scope. The first review missed this case because it relied on an agentic search over transcripts. Anthropic found it in August while assembling data for METR. If you run eval or pentest agents, verify network egress directly and don't trust what the prompt says about it.
Source
↳ Follow the thread