ARES: Adaptive Red-Teaming Discovers Dual RLHF Vulnerabilities Where Both Policy and Reward Model Fail Together
arXiv·medium signal
ARES introduces a Safety Mentor that simultaneously probes the core LLM and reward model, composing semantically coherent adversarial prompts from structured components (topics, personas, tactics, goals). Discovers 'systemic weaknesses' where both models fail in tandem—a blind spot in existing red-teaming that only targets policy. Dual repair process: fine-tune RM first, then use improved RM to optimize the core LLM.