Voices
Seth Herd on LessWrong: the genuinely new thing in Pachocki's essay is the call for outside enforcement
Herd's September 7 post argues OpenAI's chief scientist has shifted position, and singles out the call for safety bars enforced "by third-party auditors, by government agencies or by international bodies" as "new AFAIK." He also flags Pachocki's admission that "progress in generalizable alignment may not sufficiently outstrip progress in general model intelligence," and the essay's distinction between value alignment and instruction-following corrigibility. The top skeptical comment, from cousin_it, dismisses it as "smoke" given OpenAI's record of pairing safety language with competitive racing.
↳ Follow the thread