Learning to Attack and Defend: Adaptive Red Teaming of Language Models via GRPO
arXiv·medium signal
Frames AI red teaming as a co-training problem using GRPO reinforcement learning, where attacker and defender adapt against each other to discover novel attacks and produce more robust models. Moves red teaming beyond static, one-shot adversarial sets. Useful template for builders who want a continuously-evolving safety evaluation loop rather than a fixed jailbreak suite.