Dario Amodei's 'We Must Pace the Frontier' Commits Anthropic to Embedded Third-Party Evaluators, and Altman Matched It Within Hours
Amodei published the essay on September 12 arguing frontier labs should deliberately slow capability gains, and committed Anthropic unilaterally to giving outside evaluation teams such as METR ongoing 'employee-like access' to training pipelines, not just finished models. He proposes capability checkpoints where a threshold like 'the model is capable of escaping or defeating most common sandboxing methods' would require certified alignment properties before release, and warns an AI botnet could take over the internet within 6-12 months if capabilities advance unchecked. Sam Altman said OpenAI agrees and will match the first commitment; the essay took the top HN slot at 673 points and 940 comments and spawned at least four separate front-page rebuttals the same day.
↳ Follow the thread