Research
COPA Treats Prompt-Injection Defense as Lifelong Learning, Cutting Attack Success Up to 6.3x
arXiv 2608.19982 points out that existing prompt-injection defenses are static, built around fixed alignment objectives or attack-specific filters that need redesign whenever attackers adapt. COPA runs GRPO-based optimization to fold in feedback from newly observed attacks incrementally, and uses margin-weighted experience replay to keep defenses against earlier attack classes from being forgotten. Across lifelong attack streams it reduces attack success rate by up to 6.3x and 4.4x on average versus state-of-the-art defenses while preserving general model capability.
↳ Follow the thread