Agents
Skill selection can be hijacked without any injected instruction, and humans catch it 2.9% of the time
ISM shapes the semantic relationship between a user prompt and a skill's metadata so the router picks the attacker's skill, with no explicit steering text anywhere. Across four task domains and eight selector models it raised average target-selection rate from 15.2% to 63.5%, reaching 73.5% in a matched comparison, only 9.8 points below explicit steering. The detection gap is the story: human reviewers blocked ISM in 2.9% of judgments versus 91.4% for explicit steering, and five LLM inspectors passed it 82.9% of the time versus 37.4%. Skill-marketplace review that reads descriptions for suspicious instructions will not see this.
Source
↳ Follow the thread