Research
Raising a Stated Failure Probability From 10% to 70% Changes Whether Frontier Models Check the Evidence by at Most 21 Points
arXiv 2609.17865 (15 Sep 2026) evaluates an earlier decision point than usual safety benchmarks: whether a model chooses to acquire safety-relevant evidence before acting. On the SAFE benchmark across GPT-5.5, o3, Claude Opus 4.8 and Claude Sonnet 4.6, inspection policies differ sharply, with Opus inspecting nearly by default and o3 the most skip-heavy and threshold-sensitive. Inspection rises strongly with severity and falls with retrieval cost, but stated probability barely moves it, and a cost-obligation decomposition shows avoidance is driven by retrieval friction and explicit threats to the deployment payoff rather than by the duties that knowing would create.
↳ Follow the thread