Skills
Treat skill promotion as a security boundary: 3 poisoned records in a 30-record batch embed attacker behavior in 91% of trials
PoisonedEvolution attacks the promotion step where self-evolving agents distill trajectories into persistent skills — the moment untrusted experience becomes trusted instruction. A black-box attacker who can only contribute bounded evidence embedded target behaviors in 546/600 trials (91.0%) across six LLM evolvers in SkillClaw at 10% attacker support, and 369/600 (61.5%) on the structurally different Trace2Skill pipeline. Ablations pin success on recurring support, causal framing, and domain-aligned encoding; three consistent records in a 30-record batch suffice, while one is much weaker. If you auto-promote skills from traces, gate on provenance diversity, not just repetition.
↳ Follow the thread