Co-evolving the harness alongside the policy cut AgentDojo attack success 3x while raising benign utility
SafeEvolve (arXiv 2609.02786, 2026-09-02) argues that safety alignment done only at the policy level or only through external harness updates leaves runtime control disconnected from intrinsic safety. It runs a continual loop where completed on-policy trajectories become safety evidence: on the harness side, bounded component-level updates to the safety prompt and hierarchical skills, producing auditable and reversible harness artifacts; on the policy side, harness-use SFT to teach the model to actually use those artifacts, then harness-augmented RL with verifier-decomposed rewards. For Qwen3.5-4B it delivered a 3x attack-success-rate reduction on AgentDojo while benign utility rose from 59.79% to 61.86%.
↳ Follow the thread