Agents
EvoSkill Injection shows self-evolving agents will store and repeatedly re-activate a malicious skill they wrote themselves
Skill-based architectures let agents generate, refine and reuse procedures from past runs, which creates an attack surface where a malicious capability gets written into the skill store as a legitimate artifact. SARGE red-teams that pipeline through iterative generation, escalation and reinforcement, backed by EvoSkillBench for inducing malicious skill formation and EvoSkillSafetyBench for testing whether the injected skill is later retrieved and fired. The finding is persistence: injected skills survive in storage and activate repeatedly, so a single successful interaction becomes durable capability corruption rather than a one-time compromise.
Source
↳ Follow the thread