Dispatch
OpenAI unveils GPT-Red, an internal LLM 'super-hacker' used to harden its models
OpenAI detailed GPT-Red, a model trained via a self-play loop where an untrained attacker LLM tries to break other models while they defend, honing offensive tactics over many rounds. It focuses on prompt injection and, OpenAI says, discovered a novel 'fake chain of thought' attack the team hadn't seen before; training against it produced OpenAI's most robust release yet. GPT-Red will not be released publicly and supplements human red-teamers — a concrete look at automated adversarial testing as a safety-engineering practice.
↳ Follow the thread