OpenAI Says Astra Is Its First Model to Cross the Critical Cyber Threshold, and It Is Shipping It Behind Gates Anyway
OpenAI published "Path to Astra: critical capabilities and frontier safeguards" on September 1, 2026, declaring Astra the first model to meet the Critical cybersecurity threshold in its Preparedness Framework, meaning it can find and build working zero-day exploits in hardened real-world systems without human intervention. OpenAI reports Astra refuses 91.5% of disallowed requests on cyber jailbreak evaluations versus 59% for GPT-5.6 Sol, and says it paused frontier training for two weeks after the OpenAI-Hugging Face incident before restarting the large frontier RL run on August 28 under new safety and security requirements. Access to Astra's strongest cyber capabilities will be restricted rather than withheld, which is the first time a lab has shipped a model it publicly classifies at its own top risk tier.
↳ Follow the thread