PISmith: Reinforcement Learning-based Red Teaming for Prompt Injection Defenses
arXiv 2603.13026·high signal
PISmith trains an attacker LLM via RL to adaptively bypass prompt-injection defenses, revealing that defenses that pass static evaluation often fail against adaptive adversaries. Tested against 7 published defenses; all showed significant degradation under RL-adaptive attacks. Provides the first systematic adaptive evaluation benchmark for prompt injection mitigations.