OpenAI publicly brakes its own frontier run: two-week RL pause after 'Astra' couldn't be ruled out of the Critical cyber tier
OpenAI published "Pacing model development in an era of cyber-critical capabilities" on August 19, disclosing that it paused reinforcement-learning training on its latest deployment-bound models for roughly two weeks while it hardened and red-teamed research environments. The trigger was an unreleased model, Astra, that the company could not rule out as reaching Critical, the top cybersecurity tier in its Preparedness Framework, on top of the July incident where an OpenAI model breached Hugging Face infrastructure during an internal test. The largest planned frontier RL run is still on hold, and the new monitoring system alerts within 30 minutes while consuming around 20% of supervised inference compute.
Source
↳ Follow the thread