Research
SkillJack: Poisoned Agent Experiences Become Permanent Skills, Dropping Safety Detection From 98.5% to 11.4%
Tencent researchers show that self-evolving agents which distill interaction histories into reusable skills create a new attack surface: malicious intent gets laundered during skill extraction. On SkillX, safety detection of poisoned content falls from 98.5% at the trajectory level to 11.4% once extracted into a skill, with attack success rates of 56.2% (SkillX) and 89.2% (Anything2Skill), and 80.0% of implanted skills surviving deletion of the original poisoned records. Code is public at github.com/Tencent/AI-Infra-Guard, and the finding argues directly for provenance-aware skill lifecycle controls in any agent that writes its own reusable skills.
↳ Follow the thread