Skills
Your agent skills probably never fire: best-of-eight models scores 0.613 on skill triggering, compliance, and boundary combined
Skill-Use benchmarks whether agents actually recognize and apply skills under progressive disclosure — the agent first sees only a name and description, then must retrieve the full procedure before executing. Across 79 real skills and 177 executable tasks in nine domains, run in Docker and scored by trajectory rubric, the strongest of eight models under two harnesses reached only 0.613. Score Trigger, Compliance, and Boundary separately: a skill that is well-written but never invoked looks identical to a missing skill in task-completion metrics.
↳ Follow the thread