83 smart-contract audit skills tested on EVMBench: the model mattered more than the harness, and whether the skill triggered at all was the main bottleneck
arXiv 2609.29454: Demystifying Agent Skills for Smart Contract Auditing·low signal
The authors collected 83 audit skills and ran them on EVMBench across seven agent-model configurations. The biggest gain was 22.8% in detection score and 43.2% in captured award, both with Codex/GPT-5.5. Skills that triggered kept a shared six-stage audit workflow, but many never loaded. Before you tune a skill's body, confirm it actually triggers on the prompts you care about, because a strong skill that never loads adds nothing.