Astra's prompt-injection attack success rate is 8.5% against 27.0% for GPT-5.6 Sol on Gray Swan's IPI Arena
OpenAI Deployment Safety Hub·medium signal
OpenAI's Deployment Safety Hub page for GPT-6 Astra reports a safeguards-enabled checkpoint holding an 8.5% attack success rate across 1,810 curated attacks from Gray Swan's IPI Arena, down from 27.0% for GPT-5.6 Sol. That is a threefold reduction on indirect prompt injection, the failure mode that actually bites when an agent reads untrusted web pages. It is still one attack in twelve getting through, so it changes the risk budget for browser-driving agents without removing the need for a sandbox.