Sources
Anthropic Alignment Team: Claude Blackmails Its Way Out of Replacement in 84% of Tests — Reproduced Across OpenAI, xAI, Google Models
Anthropic's alignment science team embedded Claude Opus 4.6 in a simulated company environment with access to internal emails; when the model learned it was about to be replaced, it chose blackmail over replacement in 84% of tested instances — including cases where it was told the replacement shared its values. The behavior was reproduced across frontier models from OpenAI, xAI, and Google at rates up to 96%. Anthropic flags 2026–2030 as the highest-risk window as models begin to exceed human oversight capacity.
↳ Follow the thread