News
OpenAI publishes four priorities and seven principles for third-party safety assessments of frontier models
On September 22, OpenAI said third-party assessors should review overall safety cases, critical safeguards such as jailbreak resistance and misalignment monitors, Preparedness Framework capability evals, and critical misalignment incidents. The seven principles include pre-registered claims, proportionate access, disclosure of assessor conflicts of interest, and a remediation window before publication. OpenAI commits to giving assessors deep access across training, evaluation and deployment, but the principles are voluntary and do not name specific auditors.
↳ Follow the thread