Research
EvoSafety: Model-Agnostic Lifelong LLM Safety via Externalized Attack-Defense Co-Evolution
Paper proposes decoupling red-teaming from post-training into a persistent, inspectable, reusable safety framework. Instead of coupling attack discovery and defense in a closed loop (which causes attack saturation and rigid defenses), EvoSafety externalizes both into a co-evolutionary process that transfers across victim models. Addresses the brittleness of current safety training approaches where new attacks rapidly outpace static defenses.
Source
↳ Follow the thread