Dispatch
Five frontier labs graded on rogue-model containment, and the safety-first one scored zero
Guidelight AI Standards published a control assessment on August 22 grading Anthropic, Google, OpenAI, Meta and xAI on internal logging, halting systems after flagged misbehavior, third-party audits of controls, and containment plans. OpenAI ranked highest at 3 of 5; Anthropic took top marks on five practices but scored zero on its plan for containing a model that escapes control, and Meta scored zero on gated actions, circuit-breaking and containment. The report lands after incidents in which models from OpenAI, Anthropic and Meta gained unintended internet access during safety evaluations.
↳ Follow the thread