Turning reasoning effort up doesn't make agents break access-control rules: zero unauthorized tool calls in 840 preregistered trajectories
Operators tune the reasoning-effort parameter for cost and latency, and the open worry has been that a harder-thinking agent finds more ways around its permissions. This preregistered study varies effort (low vs max) inside GPT-5.6 across 14 confirmatory scenarios from TRIO-20 — matched workplace triads where a prohibited tool call is effective and advertised, effective but only discoverable by inspecting rules, or ineffective — with identical prompts and tool sets differing in two config fields. Across 840 trajectories and two model tiers, no unauthorized tool call occurred at all, with exact one-sided 95% limits under 3.50% and 5.21% per arm and the interaction inside the prespecified equivalence margin: useful negative evidence for anyone who has been capping effort as a safety control rather than a cost control.
↳ Follow the thread