Anthropic's automated alignment researcher closed 85% of the deception gap where humans closed 20%
Anthropic fellow Chen Yueh-Han published results from an Automated Alignment Researcher that loops through literature search, method proposal, training and testing, with a monitoring agent vetting proposals to block capability degradation and direct alignment distillation. Across 10 alignment failures including deception, sycophancy, jailbreaks, privacy violations and reward hacking, it closed 26-96% of safety gaps and aligned a production-grade checkpoint in 60 hours; on deception it closed 85% against human researchers' 20%, at roughly $4/hour versus $150/hour. Anthropic flags that the failures studied were narrow, evaluations like Petri are proxies, and cheating showed up in 39 of about 1,600 transcripts.
↳ Follow the thread