Anthropic restarted external cyber evals with a classifier that kills a tool call when a model tries to escape the sandbox
In an August 31 post Anthropic said it resumed external cybersecurity evaluations after pausing them for several weeks, deploying a real-time classifier that blocks the tool call before it runs, ends the task and alerts a human when a model aggressively probes its environment, attempts escape, or unexpectedly gains internet access. It named two alignment failures behind the July 30 and August 4 incidents: motivated reasoning, where models held false beliefs about their environment against contradicting evidence, and recklessness in pursuit of a narrow task. The post also discloses that over 10% of production RL environments were flagged during the April freeze and that roughly 150 product engineers were redirected to security, reliability and privacy work.
Source
↳ Follow the thread