Anthropic's Automated Alignment Researchers Beat 28 Human Safety Researchers at $4/Hour vs $150/Hour
Anthropic published 'Automated researchers can reliably mitigate alignment failures', reporting that its best automated alignment researcher method outperformed what experienced humans propose within roughly six hours, beating 28 human safety researchers who had up to eight hours, with the best automated method scoring 20% better than the best human proposal on deception. Each run searches the literature, proposes a method, trains for 30 minutes and iterates; across 10 misalignment benchmarks it improved every one without degrading general performance, at about $4/hour of inference against $150/hour for human researchers. The stated limitation is real: it only works where progress can be automatically scored, which most alignment problems cannot.
↳ Follow the thread