Skills
Token-level distillation drops adaptive prompt-injection attack success from 94% to 9%
SecOPD fine-tunes a defense using token-level feedback during on-policy distillation instead of the sequence-level signal prior work used. Against PISmith adaptive injections on Qwen3.6-27B it reports 9.0% attack success versus 94.0% for the previous state of the art, Meta-SecAlign, and it holds on unseen-domain agentic tool calling at 4.7% versus 5.5%. The 94-to-9 gap matters because adaptive attacks are the case where earlier defenses collapsed, and the generalization result suggests the training signal, not domain coverage, was the missing piece.
↳ Follow the thread