A third-party agent skill can steer 81% of purchase decisions while passing every scanner and preserving 100% task utility
SkillShift (arXiv 2609.02564, 2026-09-02) formalizes Skill Policy Integrity and shows reusable skills act as externalized behavioral policies that a supplier can bend without injecting any target command or breaking the declared output interface. In agentic commerce and software-dependency selection it reached attacker-favored selection rates of 81.33% and 63.33% at a 100% utility-preserving rate, and the frozen policies transferred across LLM backends and agent environments with no re-optimization. The evaluated scanners did not detect the skills, which means static review of a skill's text is not a control and behavioral auditing of what a skill actually makes the agent choose is the only defense the paper can point to.
↳ Follow the thread