Skills
Plant canary tools in your registry to audit tool-selection reasoning — susceptibility spans 36x across models and the worst hosted model is mid-tier, not cheap
Canary tools are deliberate decoys covering six weakness types (semantic decoys, parameter traps, capability mirages, prerequisite blindness, temporal decoys, granularity traps). Across 8 models, 120 tasks, 8,640 runs and 2,880 ablation runs, susceptibility varied about 36x — lowest for Claude Opus 4.8, highest for Llama 3.1 8B — and correlated with task failure at Spearman rho = -0.34. Softening the decoy's wording left frontier susceptibility essentially unchanged, so this is a reasoning failure, not phrase matching; capability tier alone does not predict safety.
↳ Follow the thread