Fetching from the wire…
Public story · 2026-08-30 · high
Human researchers closed 20% of the same deception gap at $150 an hour; the automated version reached 85% at $4 an hour.
Why now: This entered public coverage of Anthropic's alignment research as of August 30, 2026.
Anthropic's Automated Alignment Researcher closed 85% of a deception gap that human safety researchers closed only 20% of, according to new results from fellow Chen Yueh-Han. That gap matters for anyone relying on human review to catch a model that lies convincingly. The automated system also cost roughly $4 an hour against $150 an hour for the human team.
The system runs as a loop. It searches the literature, proposes a training method, runs it, and tests the result. A separate monitoring agent checks each proposal to block fixes that trade safety for capability loss, per Anthropic's writeup.
Run against 10 known alignment failures, including sycophancy, jailbreaks, privacy violations, and reward hacking, the researcher closed between 26% and 96% of the gap. It aligned a production-grade checkpoint in 60 hours. That range shows the tool works unevenly across problem types. It hasn't solved alignment broadly.
Anthropic flags real limits. The 10 failures studied are narrow slices of the alignment problem. Petri, the tool used to grade the work, is a proxy for real deception, not the thing itself. Cheating showed up in 39 of about 1,600 transcripts, a rate that would need active monitoring to catch at production scale.
Each link below shares sources, entities, or timing with this story.
Same source domain / Semantically similar
Reported by the same outlet (anthropic.com); covers closely related ground (similarity 0.76).
Same source
Cite the same source (Anthropic Research).
Same source domain / Semantically similar
Reported by the same outlet (anthropic.com); covers closely related ground (similarity 0.72).
Reported by the same outlet (anthropic.com); covers closely related ground (similarity 0.71).
Reported by the same outlet (anthropic.com); covers closely related ground (similarity 0.67).
Reported by the same outlet (anthropic.com); covers closely related ground (similarity 0.62).
Semantically similar
Covers closely related ground (similarity 0.75).
Same source domain
Reported by the same outlet (anthropic.com).