OpenAI's GPT-Red Is 'the Single-Largest LLM Safety Training Run Ever Documented' and Beats Human Red-Teamers at Finding Prompt Injections
An 18-author OpenAI paper submitted July 28 (Eric Wallace, Milad Nasr, Kai Xiao and others) describes a red-teaming agent trained by self-play against a population of defenders that 'reliably breaks our past models up to GPT-5.5' and 'finds more successful attacks than human red-teamers,' generalizing to held-out environments, different defender models, and different harnesses. The compute claim is the headline: the authors call it the single-largest documented LLM safety training run, comparable to their largest RL post-training efforts, and say it was used to adversarially train GPT-5.6 into their most injection-robust model to date. It sat at 4 upvotes on Hugging Face Daily Papers — badly underexposed for a paper that says defensive robustness is now a function of how much RL compute you can throw at attacking yourself, which is not a budget indie builders have.
↳ Follow the thread